BuildTube Preview About
Back to the feed

A 9B vision model built for spatial reasoning and robot planning

ZDTaichu5.0-9B is a 9-billion-parameter vision-language model, meaning it reads images, video and text together.

It pairs a Qwen3.5 language backbone with a separate vision encoder (the part that turns pictures into something the language model can read). The creator positions it for spatial reasoning, which means judging where things are and how far apart they are, including across several camera views and in 3D scenes. It also handles general image questions, document and chart reading, and multi-step tool use, and the creator aims it at robotics research where a model must plan actions in physical space.

You could use it to…

  • Ask a vision model to judge distance and depth in a photo
  • Feed it multiple camera views of the same 3D scene
  • Use it for early robotics or embodied-AI research

Can I use this?

Needs a GPU

You'll need A GPU and the transformers library · Setup needed

Worth knowing Performance comparisons against other models are the creator's own reported benchmark results at a stated model-size class, not an independent evaluation, and using it for real robotics or embodied-AI applications needs further integration work beyond the base model. Licence not stated in the excerpt available. BuildTube has not run or verified this model.

Why it's here

Getting attention on Hugging Face

  • 2,662 people have liked it on Hugging Face.
  • It was downloaded 12,395 times in the last 30 days.

Numbers from the snapshot taken 3 October 2026; not refreshed since.

Behind it

See the code on Hugging Face

BuildTube has not run or verified this project. Everything above is written from what the creator published.

Made this?

Back to the feed