BuildTube Preview About
Back to the feed

Get the same answers from a 27B model with half the thinking

A compressed, ready-to-run version of Swift-Qwen3.8-27B, which is Qwen3.8-27B adjusted to reach its answers with far less internal deliberation.

Reasoning models work by writing out a long chain of thought before answering, and that chain is what you wait for and pay for. The creator reports this version using 58.3 per cent fewer thinking tokens for under 1 per cent difference in results, giving roughly a 1.95 times speed-up on several tasks. It ships as GGUF files for llama.cpp, so the shorter reasoning and the smaller memory footprint come together.

Can I use this?

Runs with llama.cpp on your computer

You'll need llama.cpp and a capable machine · Setup needed

Worth knowing The licence is the thing to check first: this is published under a custom licence rather than a standard open one, with commercial terms offered separately, so read it before building anything on top. Every benchmark in the card is the creator’s own, and the shortened reasoning does cost accuracy in places — their own table shows a drop of several points on one mathematics benchmark, alongside a gain on one coding benchmark. The quantized comparisons come from the source model’s card and were measured on different builds than these GGUF files. BuildTube has not run or verified this model.

Why it's here

Not trending this week

Last seen on the Hugging Face trending list on 20 September 2026. It stays on BuildTube so you can still find it.

  • 453 people have liked it on Hugging Face.
  • It was downloaded 382,709 times in the last 30 days.

Numbers checked on 6 October 2026; not refreshed since.

Behind it

See the code on Hugging Face

BuildTube has not run or verified this project. Everything above is written from what the creator published.

Made this?

Back to the feed