MiniMax released its M2 series of mixture-of-experts language models, designed around what the company calls "mini activations unleashing maximum real-world intelligence." The flagship M2 model contains 229.9 billion total parameters with only 9.8 billion activated per token.
The Chinese AI lab built the M2 series specifically for agentic deployment across coding and collaborative work environments. The models rest on three core components: agent-driven data pipelines, a reinforcement learning system called Forge, and specialized inference optimization.
The data pipeline produces what MiniMax describes as "large-scale, verifiable trajectories" across agentic coding and collaborative work scenarios. Each trajectory is grounded in an executable workspace with artifact-aligned rewards.
Forge represents MiniMax's scalable agent-native RL system that adapts to long-horizon agent trajectories. The system includes windowed-FIFO scheduling, prefix-tree merging, and inference optimization designed for extended agent interactions.
Agent-First Architecture
The M2 series marks a departure from traditional language model design by prioritizing agentic capabilities from the ground up. MiniMax trained the models on trajectories that simulate real-world agent behavior rather than standard text completion tasks.
The mixture-of-experts architecture allows the model to maintain high capability while keeping computational costs manageable through selective parameter activation. Only 4.3% of the model's total parameters activate for any given token.
MiniMax joins other Chinese AI labs like Moonshot AI in releasing large-scale models designed for specific use cases. The company previously gained attention for its Hailuo video generation model and Talkie companion products.
The research paper details the technical architecture but does not specify commercial availability or pricing for the M2 series. MiniMax has not announced partnerships with cloud providers or enterprise customers for the new models.
The M2 release comes as Chinese AI companies increasingly focus on specialized model architectures rather than competing directly on parameter count with frontier labs like OpenAI and Anthropic.
💬 Discussion
Sign in to join the discussion.
Sign in →No comments yet — be the first.