PrismML released Bonsai 2 27B on Thursday, a reasoning language model compressed to 5.9 GB - small enough to run on a PC or high-end smartphone. The model retains 98% of the benchmark scores of Alibaba's Qwen3.8 27B while shrinking its memory footprint by roughly 9x to 10x.
The startup, founded by Caltech researchers and led by compression expert and professor Babak Hassibi, is betting that capable reasoning models do not need to be large. It has already seen over 11 million downloads of its first Bonsai model, released in March, with another 2.6 million downloads of its even smaller models.
How ternary weights shrink the model
PrismML's compression technique targets the weights a model learns during training. Standard models store each weight using 16 bits. The company's approach reduces each weight to one of three values: +1, -1, or 0. Storing far smaller values for each weight cuts the model's size dramatically.
This method falls under the umbrella of Generative AI and LLM research that seeks to make advanced models practical on consumer hardware. Ion Stoica, co-founder of Databricks and director of Berkeley's Sky Computing Lab, serves as an adviser to PrismML. Investors include Khosla Ventures, Cerberus Capital, and Caltech.
Performance parity inches closer
The first Bonsai model matched 95% of Qwen's aggregate benchmark scores. Bonsai 2 raises that to 98%. Hassibi said compression will likely always have some impact, but the gap is narrowing. "The next models that we will release, hopefully in the next couple of months, will be in the several-hundred-billion-parameter range, and I expect it will be easier to retain the intelligence there," he said.
He added that larger models offer more room to compress without losing capability. "There is more room to be able to compress them without losing the intelligence. So I would just say, as a general trend, for larger models, it's easier to get to 100%."
Benchmark parity itself is an academic concern. Uncompressed models are not perfectly accurate, and benchmarks do not perfectly reflect real-world tasks. A 2% degradation is unlikely to change how a model performs in practice. The software harness around the model also matters significantly for accuracy.
On-device intelligence without the cloud
Stoica sees the technology unlocking private, free AI that runs locally. "You are going to have intelligence at your fingertips, and it's going to be free because it's going to run on the device you already bought. It's also going to be private, because you're not going to send it to the cloud," he said.
PrismML is not alone in LLM compression. Multiverse Computing, founded by a professor from Spain's Donostia International Physics Center, is another competitor that has raised significant funding. Hassibi maintains that PrismML's approach is distinct because its models lose virtually no performance compared with the originals.
The company is reportedly in talks with Apple, though Hassibi declined to comment on that.
Why this matters for product development and research teams
For teams building AI features, Bonsai 2 demonstrates that a 9x-compressed model can run locally while matching 98% of the original's benchmarks. That shifts the cost and privacy equation: on-device inference removes cloud compute bills and keeps user data local. Developers evaluating model options should track whether PrismML's next release - targeting models with hundreds of billions of parameters - closes the remaining 2% gap, as the company predicts. The compression technique also has implications for AI for Science & Research, where shrinking large models without meaningful accuracy loss could make experimental AI tools portable across lab equipment and field devices.
Your membership also unlocks: