The team behind Moonshot AI's Kimi chatbot has published research on Kimi Linear, a new attention architecture that claims to outperform full attention models across multiple scenarios while dramatically improving efficiency.
The architecture centers on Kimi Delta Attention (KDA), which extends Gated DeltaNet with what the researchers call a "finer-grained gating mechanism." This allows more effective use of finite-state RNN memory compared to existing linear attention methods.
The research team pretrained a model with 3 billion activated parameters and 48 billion total parameters. The architecture combines KDA with Multi-Head Latent Attention (MLA) in a layerwise hybrid approach.
Performance and efficiency gains
In testing, Kimi Linear outperformed full MLA models across all evaluated tasks when using identical training recipes. The architecture reduced KV cache usage by up to 75% while achieving up to 6 times faster decoding throughput for 1 million token contexts.
The efficiency comes from a specialized variant of Diagonal-Plus-Low-Rank (DPLR) transition matrices. This approach substantially reduces computation compared to general DPLR formulations while maintaining consistency with classical delta rule principles.
The researchers tested the architecture across short-context, long-context, and reinforcement learning scaling scenarios. Performance improvements held across all three regimes, suggesting broad applicability.
Moonshot AI has open-sourced the KDA kernel and vLLM implementations alongside pre-trained and instruction-tuned model checkpoints. The company positions Kimi Linear as a "drop-in replacement" for full attention architectures.
The research paper lists 59 authors from the Kimi team, indicating significant internal investment in the architecture's development. The work builds on Moonshot AI's existing focus on long-context AI applications through its Kimi chatbot platform.
💬 Discussion
Sign in to join the discussion.
Sign in →No comments yet — be the first.