
Picture published by the creator on Hugging Face.
Score and rank AI agent actions instead of generating text
CLM-v0.1-8B is an 8-billion-parameter model with an unusual job: instead of writing text, it scores the choices you hand it — which action an AI agent should take next.
It bolts two small projection layers onto a frozen Qwen3-8B encoder (the text-reading part stays fixed while the new layers learn to compare states and actions). States and actions are encoded separately, so action scores can be cached and reused; the creator reports it runs up to 13 times faster than a comparable setup when ranking around a thousand candidates. It is Apache 2.0 licensed and runs through vLLM. The creator reports fine-tuned variants set records on the DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%) agent benchmarks, while noting those figures come from fine-tuned versions, not this base checkpoint.
Can I use this?
You'll need Python with vLLM and a GPU · Setup needed
Worth knowing The weights are Apache 2.0, but the model is locked to Qwen3-8B embeddings, generates no text of its own, and its scores are only meaningful relative to the candidates you supply; the headline benchmark figures are from fine-tuned variants, not this checkpoint run as-is. BuildTube has not run or verified this model.
Why it's here
- 584 people have liked it on Hugging Face.
- It was downloaded 2,720 times in the last 30 days.
Numbers from the snapshot taken 1 October 2026; not refreshed since.
Behind it
See the code on Hugging FaceBuildTube has not run or verified this project. Everything above is written from what the creator published.
Made this?