BuildTube Preview About
Back to the feed

Run a 35B-class chat model in about 3 GB of memory

A large language model that keeps most of itself on storage and pulls in only the pieces it needs for each word, so it runs in far less memory than its size suggests.

The creator reports under 3 GB of active memory and 15 words per second on a laptop.

Can I use this?

Setup details on the source page

You'll need Apple Silicon and the edge0 framework · Setup needed

Worth knowing The creator labels this an early preview. Every memory and speed figure is its own, measured with its own streaming framework, and the model needs that framework rather than running in the usual tools. Streaming from storage means disk speed matters as much as memory. Released under Apache 2.0. BuildTube has not run or verified this model.

Why it's here

Not trending this week

Last seen on the Hugging Face trending list on 18 September 2026. It stays on BuildTube so you can still find it.

  • 3,566 people have liked it on Hugging Face.
  • It was downloaded 81,263 times in the last 30 days.

Numbers checked on 6 October 2026; not refreshed since.

Behind it

See the code on Hugging Face

BuildTube has not run or verified this project. Everything above is written from what the creator published.

Made this?

Back to the feed