xAI launched Grok 4.5, its most advanced AI model designed for coding, agentic tasks, and knowledge work.
The model achieves a 62% score on DeepSWE 1.0, placing it third behind Fable's 66.1% and GPT 5.5's 64.31%. On SWE Marathon, Grok 4.5 leads with a 29% resolution rate, outperforming Opus 4.8's 26% and Fable's 24%.
Grok 4.5 was trained on tens of thousands of NVIDIA GB300 GPUs using datasets spanning coding, science, engineering, and mathematics. The training incorporated reinforcement learning across hundreds of thousands of tasks, with automated grading systems running asynchronous rollouts for hours.
xAI partnered with Cursor during development, and the model is now available across Cursor's plans. The collaboration focused on real-world engineering tasks and multi-step software development challenges.
Token efficiency breakthrough
The model delivers 4.2x better token efficiency than Opus 4.8, using an average of 15,954 output tokens per SWE Bench Pro task compared to Opus 4.8's 67,020 tokens. This efficiency translates to faster results at lower costs.
Grok 4.5 serves responses at 80 tokens per second, matching flash model speeds while maintaining higher reasoning capabilities. The model can build complete applications from single prompts, including complex three.js simulations and multi-sheet Excel models.
Pricing starts at $2 per million input tokens and $6 per million output tokens. xAI positions this as competitive given the model's token efficiency advantages over comparable systems.
The model integrates with Microsoft Office applications, building PowerPoint presentations with native shapes and complex diagrams. In Excel, it creates multi-sheet formulas while leaving reference notes for future use.
Grok 4.5 launches today through the xAI API console and Grok Build platform. The company offers free usage for a limited period to encourage adoption among developers and enterprise users.
💬 Discussion
Sign in to join the discussion.
Sign in →No comments yet — be the first.