AI infrastructure startup Wafer achieved 952 tokens per second running Moonshot AI's massive Kimi K3 model on AMD's MI355X GPUs, demonstrating superior performance per dollar compared to Nvidia's flagship chips.
The 2.8 trillion parameter Kimi K3 model requires over 1.5TB of VRAM before allocating memory for its 1 million token context window. That exceeds the capacity of even Nvidia's B200 nodes, forcing deployments onto expensive B300 systems or dual-node B200 configurations.
Wafer's MI355X deployment delivered 118 tokens per second for single-stream inference and 952 tokens per second aggregate throughput. The company compared this against a 16-GPU B200 setup spanning two nodes, which managed just 498 tokens per second total.
Performance and cost breakdown
At $2.50 per GPU hour, the MI355X achieved 48 tokens per second per dollar. Nvidia's B300 managed 33 tokens per second per dollar at $6.00 per GPU hour, while the B200 delivered just 7 tokens per second per dollar at $4.25 per hour.
The MI355X's 288GB of VRAM per GPU matches the B300's capacity but costs approximately 2.4x less. This memory advantage proves crucial for models like Kimi K3 that push hardware limits.
Wafer encountered software challenges typical of AMD deployments. The team fixed a missing function in the speculative decoding verifier and optimized prefill performance by resolving a shape mismatch in attention kernels.
"The MI355X struggles here: an identical 172k-token cold prefill took ~51s on MI355X versus ~23s on a B300," Wafer noted in their technical breakdown.
The optimizations improved single-stream performance by 2.2x and boosted peak aggregate throughput by 18%. Prefill speeds increased 2-3x after implementing the attention kernel fix.
Wafer's results suggest AMD's software ecosystem is maturing rapidly. The company noted fewer framework bugs than previous deployments and no need for custom kernels, marking progress toward day-zero model support on AMD hardware.
💬 Discussion
Sign in to join the discussion.
Sign in →No comments yet — be the first.