Moonshot AI has launched K3-256k, a more efficient version of its flagship K3 coding model that maintains the same 2.8 trillion parameters while operating with a 256K context window instead of the full 1M version.

The new model consumes roughly half the quota of the standard K3 while delivering identical coding capabilities. K3-256k supports adjustable reasoning effort levels — low, high, and max — with high set as the default for optimal performance.

Kimi Code's platform now offers four distinct model configurations across three model families. The flagship K3 remains available with its full 1M context window for Allegretto-tier subscribers and above, while K3-256k opens access to Moderato-level members.

Model lineup and pricing tiers

The company also maintains K2.8 Preview, which delivers performance close to K3 with more efficient thinking processes. A high-speed variant, K2.7 Code HighSpeed, provides 5-6x faster output at 3x quota usage for rapid development workflows.

All models support multimodal input including images, with K3 and K2.8 Preview also handling video files. The reasoning effort system allows developers to tune computational intensity based on task complexity.

Moonshot AI has structured access around membership tiers, with basic members accessing K2.8 Preview, Moderato members unlocking K3-256k, and higher tiers gaining full 1M context capabilities.

The platform supports both OpenAI and Anthropic API protocols, enabling integration with third-party coding tools like Claude Code and OpenCode. Developers can switch models mid-session, though the company recommends starting fresh sessions to optimize context caching.

Kimi Code's API endpoints are available at api.kimi.com/coding/v1 for OpenAI compatibility and api.kimi.com/coding/ for Anthropic protocol support. The service competes directly with GitHub Copilot and other AI coding assistants in the enterprise development market.