PrismML shipped Bonsai 27B, the first 27-billion-parameter AI model capable of running on smartphones. The model uses extreme quantization techniques to compress what would normally be a 54GB model into variants as small as 3.9GB.

The company released two versions: a ternary variant at 5.9GB that retains 95% of the original model's performance, and a 1-bit version at 3.9GB that maintains 90% capability while fitting within an iPhone's memory constraints.

Both models are based on Qwen 3.6 27B and support multimodal tasks including vision, reasoning, and tool calling. The compression applies across the entire network architecture — embeddings, attention layers, and language modeling head — with no higher-precision components.

Performance benchmarks show minimal capability loss

Across 15 benchmarks covering math, coding, and agentic tasks, the ternary version scored 80.5 compared to the baseline's 85.0. Math and coding performance remained nearly intact, with the model scoring 93.4 versus 95.3 on mathematical reasoning tasks.

The 1-bit variant achieved 76.1 on the same benchmark suite while occupying 2.5 times less memory than conventional 4-bit quantized versions of similar models.

PrismML measured "intelligence density" at 0.53 per gigabyte for the 1-bit model — more than 10 times the full-precision baseline and 2.7 times better than existing low-bit alternatives.

On-device inference speeds reach 163 tokens per second

The models achieve up to 163 tokens per second on NVIDIA RTX 5090 GPUs and 87 tokens per second on Apple's M5 Max chips. The company demonstrated agentic workflows running entirely on-device, including tool calling and multi-step reasoning tasks.

PrismML positioned the release as enabling "hybrid deployments" where privacy-sensitive tasks run locally while complex operations route to cloud models. The approach addresses cost accumulation in multi-step agentic workflows that can require hundreds of model calls.

Both variants ship with 262K-token context windows and support speculative decoding for additional speed improvements. The models are available under Apache 2.0 licensing with native support for Apple devices via MLX and NVIDIA GPUs through custom CUDA kernels.