One model. Very different workloads.
Coding.
Compiler development.
Rust optimization.
Agent planning.
Three.js generation.
Healthcare reasoning.
Ling-3.1-flash is clearly being designed for more than chat.
560B parameters, 25B active/token, and up to 1M-token context make the architecture especially interesting.
Meet Ling-3.1-flash: ~560B total params, ~25B active/token, up to 1M-token context.
We plan to open-source the model soon.
Across work, coding & healthcare: 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, and 65.35 on HealthBench Professional.