PrismML's Compression Breakthrough Shrinks Reasoning Models to Fit on Your Phone
The AI startup has developed a technique to compress large language models down to a fraction of their original size while retaining nearly all their performance, potentially shifting how AI runs on consumer devices.

PrismML deserves attention not for its funding haul—a $22.25 million seed round—but for the caliber of its research team and the transformative potential of its work. The AI lab is challenging a fundamental assumption in the field: that reasoning models with strong performance must necessarily be large.
The company is developing reasoning models compact enough to run on personal computers and smartphones. Speculation about talks with Apple has circulated, though CEO Babak Hassibi has not confirmed such discussions.
On Thursday, PrismML unveiled Bonsai 2 27B, the latest addition to its model family. The release compresses Qwen 27B, a popular open-source model from Alibaba, to just 5.9 GB—a reduction of 9x to 10x in memory footprint compared to the original. This size makes it viable for deployment on standard PCs and potentially on premium smartphones.
Founded by researchers from Caltech and led by Hassibi, a Caltech professor specializing in compression technologies, the startup counts Ion Stoica among its advisors. Stoica co-founded Databricks and directs Berkeley's Sky Computing Lab, an incubator for technologies and ventures including Letta and SGLang. Khosla Ventures, Cerberus Capital, and Caltech have backed the company.
While other companies, such as Multiverse Computing, are also pursuing LLM compression, Hassibi contends that PrismML's approach stands apart. The company's models retain virtually all the performance of their uncompressed counterparts. Bonsai 2 achieves 98% of Qwen's aggregate benchmark scores, an improvement from the original Bonsai released in March, which matched 95%. The first version has been downloaded over 11 million times, with PrismML's smaller models accumulating an additional 2.6 million downloads.
The gap between releases suggests steady progress in compression quality. Whether the company can eventually reach 100% benchmark parity remains uncertain. Hassibi notes that compression will likely always carry some performance cost.
In practical terms, a 2% performance loss may not meaningfully impact real-world model behavior, given that uncompressed models themselves are imperfect and benchmarks don't fully capture actual task performance. The software environment surrounding a model—what Hassibi calls the harness—also significantly influences accuracy.
PrismML achieves compression by reducing the "weights" that constitute a model's learned information. Standard weights require 16 bits each. PrismML's "ternary" weight approach reduces this to three possible values: +1, −1, or 0. Storing far fewer bits per weight dramatically shrinks the overall model size.
The startup's next phase involves scaling this compression technique to substantially larger models. According to Hassibi, "The next models that we will release, hopefully in the next couple of months, will be in the several-hundred-billion-parameter range, and I expect it will be easier to retain the intelligence there." He added that larger models offer more room for compression without intelligence loss, making it "easier to get to 100%" as model size increases.
Stoica sees significant potential in this direction. He stated that the technology enables advanced models to operate directly on user devices: "You are going to have intelligence at your fingertips, and it's going to be free because it's going to run on the device you already bought. It's also going to be private, because you're not going to send it to the cloud."


