BuildTube Preview About
Back to the feed

Run a 29B model that uses only 4B of itself per word

A model from China Telecom’s AI arm with 29 billion parameters, of which about 4 billion are used for each word, so it costs less to run than its size suggests.

It handles a 256,000-token context and is built for multi-step tasks and tool calling.

You could use it to…

  • Hold a whole large codebase in one long context
  • Run multi-step tool-calling tasks at a lower cost
  • Use a 256,000-token window without paying full-size prices

Can I use this?

Needs a GPU server

You'll need A GPU server running vLLM or SGLang · Setup needed

Worth knowing The creator says this is the first model at this scale trained entirely on Huawei’s Ascend hardware, and the training-efficiency figures are its own. Most of the detail lives in a separate repository rather than on the model page. Released under Apache 2.0. BuildTube has not run or verified this model.

Why it's here

Not trending this week

Last seen on the Hugging Face trending list on 28 September 2026. It stays on BuildTube so you can still find it.

  • 1,833 people have liked it on Hugging Face.
  • It was downloaded 50,789 times in the last 30 days.

Numbers checked on 6 October 2026; not refreshed since.

Behind it

See the code on Hugging Face

BuildTube has not run or verified this project. Everything above is written from what the creator published.

Made this?

Back to the feed