Models

PrismML's Bonsai 2 27B Shrinks AI to 5.9GB Without Sacrificing Smarts

PrismML has released Bonsai 2 27B, a compressed AI model that runs on consumer PCs and high-end phones while retaining 98.2% of the capabilities of a much larger base model.

·2 min read
PrismML launches Bonsai 2 27B, a high-intelligence AI model so small it fits on consumer hardware
PrismML launches Bonsai 2 27B, a high-intelligence AI model so small it fits on consumer hardware

PrismML Inc. unveiled Bonsai 2 27B on Thursday, marking the second iteration of its compact multimodal generative AI designed to operate on standard computers and certain premium mobile devices. The startup compressed its Qwen3.8 27B-based model using ternary quantization, reducing the original 56-gigabyte uncompressed 16-bit version down to approximately 5.9 gigabytes while preserving roughly 98.2% of its original performance.

Traditional compression methods like quantization often degrade model accuracy and reasoning capabilities. Ternary quantization works differently by simplifying the model's weights—the numerical parameters controlling how the system processes information—from 16 bits down to three bits, represented as +1, 0, and -1. This approach maintains meaningful intelligence while dramatically shrinking the memory footprint. For comparison, Qwen3.8's minimum footprint stands at 9.4 gigabytes.

Testing shows Bonsai 2 performs competitively against its larger parent model. On agentic and tool-calling benchmarks, it scored 77.6 versus Qwen3.8's 79.8—a gap of just 3 points. Coding performance across HumanEval+, LiveCodeBench v6, MBPP+ and BigCodeBench reached 81.6 compared to 82.2. Knowledge and reasoning scores across MMLU-Redux, GPQA Diamond and AA-LCR came in at 82.7 versus 81.3.

Hardware performance is notable. Running on an Nvidia GeForce GTX 5090 without additional compression, the model achieves 143 tokens per second. On Apple Inc.'s M5 Max chip, it reaches 46.8 tokens per second. The model consumes 0.714 megawatt-hours per token, making it 40% more energy-efficient than other 8B models operating at full precision.

Local AI and Privacy Benefits

Running AI models locally on personal devices eliminates the need to send inference requests to cloud servers. This approach sidesteps latency issues inherent in internet communication and prevents sensitive data from reaching third-party systems. Local execution helps organizations meet strict privacy regulations while improving security posture.

A practical hybrid approach emerges: straightforward tasks like translation, summarization and search organization can execute on-device using lightweight models, while complex work requiring deeper reasoning and long-horizon task comprehension routes to expensive cloud-based systems. For both individual users and enterprises, this strategy balances cost efficiency with capability, keeping sensitive operations private while reserving computational resources for genuinely demanding workloads.

Bonsai 2 runs on Nvidia graphics processing units through CUDA and on Apple devices—Mac, iPhone and iPad—via MLX, utilizing low-bit kernels for optimization. Model weights are available today under Apache 2.0 licenses.