Fit a 27B reasoning model in 8.6 GB on an Apple Silicon Mac
Made by Prism ML
View profile on Hugging Face (opens in a new tab)The Apple Silicon build of a heavily compressed 27-billion-parameter model, storing each weight as one of three values — minus one, zero or plus one — so the whole thing including a vision component occupies 8.6 GB instead of roughly 54 GB.
The creator reports it retaining 98.2 per cent of the full-precision model’s benchmark average across fourteen reasoning tests, and running at around 47 tokens a second on an M5 Max laptop. It keeps the 262,144-token context of the model it came from, and runs through MLX, Apple’s own machine learning framework, with custom low-bit kernels.
Can I use this?
You'll need An Apple Silicon Mac with about 9 GB free · Setup needed
Worth knowing Every compression and quality figure here is the creator’s own, measured on benchmarks they selected; heavily compressed models can differ from the original in ways a benchmark average does not reveal. The custom low-bit kernels live in the creator’s own forks of MLX rather than in the standard releases, which is a real setup cost and a dependency on those forks staying maintained. Apple Silicon only — a companion GGUF build exists for other hardware. Apache 2.0. BuildTube has not run or verified this model.
Why it's here
Not trending this week
Last seen on the Hugging Face trending list on 20 September 2026. It stays on BuildTube so you can still find it.
- 430 people have liked it on Hugging Face.
- It was downloaded 74,567 times in the last 30 days.
Numbers checked on 6 October 2026; not refreshed since.
Behind it
See the code on Hugging FaceBuildTube has not run or verified this project. Everything above is written from what the creator published.
Made this?