PrismML Shrinks Model to Bring It to Personal Devices

Serdar HocamAuthor & Editor

Artificial intelligence startup PrismML has developed a new version by compressing Alibaba's open-source model so that it can run on personal computers and smartphones.

◉ 0 views
PrismML hopes its tiny LLM will change how we all use AI | TechCrunch

AI startup PrismML has announced Bonsai 2 27B, a new compressed reasoning model that shrinks Alibaba's Qwen3.8 27B model down to just 5.9 GB.

The Era of Small Size in Artificial Intelligence

PrismML is developing new technology by arguing that high-performance large language models with reasoning capabilities do not have to be massive in size. Accordingly, the startup aims to shrink the models it prepares so that they can run on personal computers and smartphones.

The Size of the Bonsai 2 27B Model

Introduced on Thursday, the Bonsai 2 27B model reduces Alibaba's widely used Qwen3.8 27B open-source model to a size of 5.9 GB. This work achieves a 9 to 10-fold reduction in memory usage compared to the original model.

Academic-Origin Team and Investors

Founded by a group of Caltech researchers, the company is led by compression technologies expert and Caltech professor Babak Hassibi. The startup, which is advised by Databricks co-founder Ion Stoica, is backed by investors including Khosla Ventures, Cerberus Capital, and Caltech.

No Performance Loss

PrismML states that thanks to the developed compression technology, the models experience almost no performance loss compared to their original versions. While the first version released in March achieved a 95% success rate, the Bonsai 2 model matches 98% of Qwen's aggregate benchmark scores.

Ternary Weight Technology

The company achieves this success by using ternary weights, which reduce traditional 16-bit weights to three distinct values (+1, -1, or 0). This method increases efficiency, allowing large models to be stored in smaller footprints.

Future Goals and New Models

The startup's next goal in the coming months is to apply this compression technique to much larger models in the hundreds of billions of parameters range and spread the technology to broader areas.