BuildTube Preview About
Back to the feed

Try the architecture behind Qwen4 as an open preview

An experimental preview of the architecture Alibaba intends to build its next model generation on, released with open weights.

It is a mixture-of-experts model — 125 billion parameters total but only about 6 billion active on any token, plus a 51-billion-parameter n-gram embedding table — so the compute cost per word is far below what the total size suggests. It reads images as well as text, holds 262,144 tokens natively and can be extended toward a million, and it introduces a sparse attention mechanism that works on blocks of tokens rather than individual ones to cut the cost of very long contexts.

Can I use this?

Needs a model runtime on your computer

You'll need Server-class GPUs, or a hosted provider · Setup needed

Worth knowing The licence is a custom one rather than a standard open licence, so check the terms before commercial use. The creator calls this an experimental preview of an architecture under development, which is a fair warning that behaviour and support may change and that the production version with longer default context and built-in tools is a separate hosted product. All benchmark comparisons are the creator’s own. Despite the small active parameter count, all 125 billion parameters still have to be held in memory. BuildTube has not run or verified this model.

Why it's here

Not trending this week

Last seen on the Hugging Face trending list on 20 September 2026. It stays on BuildTube so you can still find it.

  • 5,962 people have liked it on Hugging Face.
  • It was downloaded 1,589,145 times in the last 30 days.

Numbers checked on 6 October 2026; not refreshed since.

Behind it

See the code on Hugging Face

BuildTube has not run or verified this project. Everything above is written from what the creator published.

Made this?

Back to the feed