AI startup PrismML is aiming to transform artificial intelligence accessibility by developing advanced compression technology for large language models. The company recently released Bonsai 2 27B, successfully shrinking a widely used open-source model down to just 5.9 GB, making it small enough to run efficiently on personal computers and high-end smartphones.
Founded by Caltech researchers and advised by industry veterans, PrismML achieves this drastic reduction—roughly ten times smaller than the original—by converting traditional 16-bit weights into ternary values of +1, -1, or 0. Remarkably, this compression retains 98% of the original model’s benchmark performance while virtually eliminating noticeable quality loss in daily tasks.
Looking ahead, the startup plans to apply its compression methodology to models featuring hundreds of billions of parameters. This advancement paves the way for secure, private, and zero-cost local AI processing directly on consumer hardware without relying on cloud infrastructure.
- Bonsai 2 27B model compressed down to 5.9 GB.
- Retains 98% of the original model’s benchmark performance.
- Utilizes ternary weights to drastically reduce memory footprint.
- Enables private, local AI execution on consumer devices.
Sources:
