Mark RadarMARK RADAR
About
EN
Sign in

PrismML Packs 27B AI Model Into 5.9 GB for Edge Devices

2 reports · First detected 2026-09-18 · Last active 2026-09-19

Founded by California Institute of Technology researchers and led by Caltech professor Babak Hassibi, PrismML is pursuing low-bit compression as an alternative to ever-larger cloud models. The approach matters because shrinking capable reasoning systems enough for PCs and premium smartphones could cut memory use, latency and power consumption while keeping private data on-device. The startup has raised a $22.25 million seed round from backers including Khosla Ventures, Cerberus Capital and Caltech.

On Sept. 17, 2026, PrismML released Ternary Bonsai 2 27B, compressing Alibaba’s open-source Qwen3.8 27B from about 54 GB at FP16 to 5.9 GB, a more than ninefold reduction. PrismML said the model scored 83.9 across 20 benchmarks, retaining 98.2% of the full-precision model’s 85.4 aggregate score. The free Apache 2.0 release supports multimodal text-and-image input and a 262,000-token context window, with reported speeds of up to 143 tokens per second on Nvidia’s GeForce RTX 5090.

All Coverage

2 original reports

The Backstory

The history behind this event
PrismML Brings 1-Bit Bonsai-27B to Local Inferencefirst seen 2026-07-28 · 1 reports · similarity 0.84

Bonsai-27B is a roughly 27-billion-parameter large language model using 1-bit weights to reduce memory and computational demands. PrismML’s version of llama.cpp is designed to run the model locally with NVIDIA CUDA acceleration, offering developers an alternative to cloud-hosted inference. The approach is significant for organizations seeking greater control over data, latency and operating costs while retaining access to generative AI capabilities.

The latest deployment guide details how to compile PrismML’s CUDA kernels, download the Bonsai-27B model weights and launch a local inference server compatible with the OpenAI API. The demonstrated workflow supports multi-turn conversations and code generation while allowing applications to reuse existing OpenAI client interfaces. The report does not specify a publication date, benchmark results, hardware configuration or deployment cost, limiting direct comparisons with other local and hosted models.

Mark Radar|MARK RADAR

If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →

All times are in Taipei time (GMT+8)