Former OpenAI CTO Mira Murati's Thinking Machines Lab released Inkling, a 975-billion-parameter mixture-of-experts foundation model with open weights for full customization.
The model activates 41 billion parameters during inference and supports context windows up to one million tokens. Inkling was pretrained from scratch on 45 trillion tokens spanning text, images, audio, and video.
Inkling processes text, image, and audio inputs natively. Its capabilities include reasoning, coding, tool use, instruction following, visual analysis, speech transcription, and long-form audio understanding.
View tweet from @huggingface
Developers can adjust Inkling's thinking effort between 0.2 and 0.99 to balance output quality, latency, and token consumption. Thinking Machines says the model matches Nemotron 3 Ultra on Terminal Bench 2.1 while using roughly one-third as many generated tokens.
Performance benchmarks
Inkling scored 77.6% on SWE-bench Verified, 97.1% on AIME 2026, 87.2% on GPQA Diamond, and 73.5% on MMMU Pro. The company positions it as a broad base for custom models rather than the highest-scoring general-purpose model.
The model operates inside coding-agent harnesses, uses changing tool schemas, and produces applications and structured artifacts through repeated refinement.
Tinker offers Inkling with 64K and 256K context options at a temporary 50% discount. The Inkling Playground provides a chat interface with integrated agentic web search, free for a limited period.
API access is available through Together, Fireworks, Modal, Databricks, and Baseten. Inference support spans SGLang, vLLM, TokenSpeed, llama.cpp, and Hugging Face Transformers.
Thinking Machines is also previewing Inkling-Small, a 276-billion-parameter mixture-of-experts model with 12 billion active parameters. It approaches or surpasses Inkling on several reasoning, instruction-following, vision, and audio tests while targeting lower-cost workloads.
The release extends Thinking Machines' Tinker customization platform and serves as the reasoning layer behind its previously previewed real-time voice and vision systems. Full weights for Inkling-Small will follow after testing completes.
💬 Discussion
Sign in to join the discussion.
Sign in →No comments yet — be the first.