
Picture published by the creator on Hugging Face.
Run Qwen3.8-27B on your own machine from a compressed file
Made by Unsloth AI
View profile on Hugging Face (opens in a new tab)Compressed builds of Qwen3.8-27B, Alibaba’s 27-billion-parameter model that reads images and video as well as text, packaged as GGUF files so they run through llama.cpp, Ollama or LM Studio on hardware that could not hold the original.
The compression is Unsloth’s own Dynamic 3.0 method, which varies how hard it squeezes different parts of the model instead of applying one setting throughout; the creator reports more than 10 per cent better top-1 accuracy at the same file size than other providers’ quantizations. The underlying model has a 262,144-token context and a thinking mode you can switch off per request.
Can I use this?
You'll need llama.cpp, Ollama or LM Studio, and a good deal of memory · Setup needed
Worth knowing These are compressed copies, and compression always costs something — the comparison the creator publishes is against other compressions, not against the original at full precision. Which file to pick is a real decision, since the repository holds several sizes trading quality against memory, and a 27-billion-parameter model is demanding even squeezed. The accuracy comparisons are Unsloth’s own measurements. Apache 2.0, the same as the original model. BuildTube has not run or verified this model.
Why it's here
Not trending this week
Last seen on the Hugging Face trending list on 20 September 2026. It stays on BuildTube so you can still find it.
- 4,887 people have liked it on Hugging Face.
- It was downloaded 6,566,855 times in the last 30 days.
Numbers checked on 6 October 2026; not refreshed since.
Behind it
See the code on Hugging FaceBuildTube has not run or verified this project. Everything above is written from what the creator published.
Made this?