Run a 29B model that uses only 4B of itself per word
Made by XingChen-AGI
View profile on Hugging Face (opens in a new tab)A model from China Telecom’s AI arm with 29 billion parameters, of which about 4 billion are used for each word, so it costs less to run than its size suggests.
It handles a 256,000-token context and is built for multi-step tasks and tool calling.
You could use it to…
- Hold a whole large codebase in one long context
- Run multi-step tool-calling tasks at a lower cost
- Use a 256,000-token window without paying full-size prices
Can I use this?
You'll need A GPU server running vLLM or SGLang · Setup needed
Worth knowing The creator says this is the first model at this scale trained entirely on Huawei’s Ascend hardware, and the training-efficiency figures are its own. Most of the detail lives in a separate repository rather than on the model page. Released under Apache 2.0. BuildTube has not run or verified this model.
Why it's here
Not trending this week
Last seen on the Hugging Face trending list on 28 September 2026. It stays on BuildTube so you can still find it.
- 1,833 people have liked it on Hugging Face.
- It was downloaded 50,789 times in the last 30 days.
Numbers checked on 6 October 2026; not refreshed since.
Behind it
See the code on Hugging FaceBuildTube has not run or verified this project. Everything above is written from what the creator published.
Made this?