Chat with a 27B model on a laptop, compressed to under 6 GB
Made by PrismML-Eng
View profile on GitHub (opens in a new tab)Bonsai Demo is the official runner for PrismML's Bonsai family of heavily compressed language models, letting you run a 27-billion-parameter class model on a Mac, a single GPU, or plain CPU.
Its default model, Bonsai 2 27B, is compressed to ternary weights (each stored as one of just three values) at 5.9 gigabytes, and the creator reports it keeps 98.2% of the accuracy a full-precision version would have. Two setup commands download the right build for your machine and start a local chat server with vision (it can read photos, screenshots and PDFs), tool calling, and an adjustable reasoning effort.
You could use it to…
- Chat with a 27B-class model on a laptop, under 6GB
- Show it a screenshot or PDF and ask about it
- Turn tool calling on with full round-trips, on-device
Can I use this?
You'll need A Mac, GPU, or capable CPU, and disk space for the model · Setup needed
Worth knowing The near-full-precision accuracy figures and benchmark comparisons are the creator's own reported numbers, and this demo requires the project's own modified llama.cpp build rather than a stock installation. Apache-2.0 licence. BuildTube has not run or verified this software.
Why it's here
Not trending this week
Last seen on the GitHub trending list on 25 September 2026. It stays on BuildTube so you can still find it.
- 3,283 people have starred it on GitHub.
Numbers checked on 6 October 2026; not refreshed since.
Behind it
See the code on GitHubBuildTube has not run or verified this project. Everything above is written from what the creator published.
Made this?