Ant Group announced the release of Ling-3.0-Flash, a next-generation native hybrid-reasoning foundational model engineered specifically for production-grade AI agent workflows. Designed to deliver rapid response capabilities, it serves as a high-speed execution node that offers a superior balance of intelligence density and cost-efficiency.
Featuring 124B total parameters with only 5.1B active parameters per token, Ling-3.0-Flash achieves remarkable performance despite its streamlined footprint. It matches or surpasses industry-leading models with two to three times its parameter scale across core benchmarks, including foundational reasoning, instruction following, and long-context processing.
Ling-3.0-Flash moves away from the traditional approach of simply scaling parameter counts. Instead, it is built from the ground up with a native hybrid-linear attention architecture. By alternating KDA (Kimi Delta Attention) and MLA layers at a 5:1 ratio, the model optimally balances long-context efficiency with robust state memory.
Upgraded KDA: Evolving from the previous Lightning Attention, KDA introduces fine-grained diagonal gating in Delta Rule state updates, allowing the model to retain critical information more precisely when processing lengthy documents and extensive codebases.
Optimized Mixture-of-Experts (MoE) Compute: The expert activation ratio per token has been compressed from 1/32 in the previous generation to 1/64, yielding a significantly higher "efficiency leverage."
Extended Context Window: The model natively supports a 256K context window and can seamlessly scale to 1M tokens.
Rather than aiming to replace ultra-large, general-purpose reasoning models, Ling-3.0-Flash is designed to complete the "planning-execution separation" paradigm in AI workflows. It serves as a cost-controllable, fast, and highly stable execution node, delegating deep planning and high-frequency execution to specialized models.
To support this, Ling-3.0-Flash has been deeply refined for real-world agent scenarios, expanding its training to over 10,000 interactive environments. It features enhanced self-correction and long-horizon planning mechanisms, enabling autonomous, end-to-end delivery in complex tasks such as coding, task decomposition, and deep multi-source research. This resolves common issues of deviation or context loss in traditional models during large-scale operations.
To ensure fast and reliable agent performance, Ant Group has paired Ling-3.0-Flash with a supporting engineering and collaboration architecture:
Reduced Latency: A cluster-level hierarchical caching system eliminates redundant computations in long conversations and multi-turn interactions, reducing Time-to-First-Token (TTFT) for long inputs by 60 per cent to over 80 per cent.
Enhanced Stability: An upgraded multi-agent collaboration architecture enables different agents to divide labor and cross-validate outputs, significantly reducing the risk of misjudgments by a single model and providing robust support for high-frequency online services.


