BuildTube Preview About
Back to the feed

Fit a 27B reasoning model in 8.6 GB on an Apple Silicon Mac

The Apple Silicon build of a heavily compressed 27-billion-parameter model, storing each weight as one of three values — minus one, zero or plus one — so the whole thing including a vision component occupies 8.6 GB instead of roughly 54 GB.

The creator reports it retaining 98.2 per cent of the full-precision model’s benchmark average across fourteen reasoning tests, and running at around 47 tokens a second on an M5 Max laptop. It keeps the 262,144-token context of the model it came from, and runs through MLX, Apple’s own machine learning framework, with custom low-bit kernels.

Can I use this?

Needs a model runtime on your computer

You'll need An Apple Silicon Mac with about 9 GB free · Setup needed

Worth knowing Every compression and quality figure here is the creator’s own, measured on benchmarks they selected; heavily compressed models can differ from the original in ways a benchmark average does not reveal. The custom low-bit kernels live in the creator’s own forks of MLX rather than in the standard releases, which is a real setup cost and a dependency on those forks staying maintained. Apple Silicon only — a companion GGUF build exists for other hardware. Apache 2.0. BuildTube has not run or verified this model.

Why it's here

Not trending this week

Last seen on the Hugging Face trending list on 20 September 2026. It stays on BuildTube so you can still find it.

  • 430 people have liked it on Hugging Face.
  • It was downloaded 74,567 times in the last 30 days.

Numbers checked on 6 October 2026; not refreshed since.

Behind it

See the code on Hugging Face

BuildTube has not run or verified this project. Everything above is written from what the creator published.

Made this?

Back to the feed