Get the same answers from a 27B model with half the thinking
Made by UkisAI
View profile on Hugging Face (opens in a new tab)A compressed, ready-to-run version of Swift-Qwen3.8-27B, which is Qwen3.8-27B adjusted to reach its answers with far less internal deliberation.
Reasoning models work by writing out a long chain of thought before answering, and that chain is what you wait for and pay for. The creator reports this version using 58.3 per cent fewer thinking tokens for under 1 per cent difference in results, giving roughly a 1.95 times speed-up on several tasks. It ships as GGUF files for llama.cpp, so the shorter reasoning and the smaller memory footprint come together.
Can I use this?
You'll need llama.cpp and a capable machine · Setup needed
Worth knowing The licence is the thing to check first: this is published under a custom licence rather than a standard open one, with commercial terms offered separately, so read it before building anything on top. Every benchmark in the card is the creator’s own, and the shortened reasoning does cost accuracy in places — their own table shows a drop of several points on one mathematics benchmark, alongside a gain on one coding benchmark. The quantized comparisons come from the source model’s card and were measured on different builds than these GGUF files. BuildTube has not run or verified this model.
Why it's here
Not trending this week
Last seen on the Hugging Face trending list on 20 September 2026. It stays on BuildTube so you can still find it.
- 453 people have liked it on Hugging Face.
- It was downloaded 382,709 times in the last 30 days.
Numbers checked on 6 October 2026; not refreshed since.
Behind it
See the code on Hugging FaceBuildTube has not run or verified this project. Everything above is written from what the creator published.
Made this?