A large model trained on coding, agents, vision and security tasks at once
Made by Xiaomi MiMo
View profile on Hugging Face (opens in a new tab)MiMo-V2.6-Pro-RL is Xiaomi's flagship model in the MiMo-V2.6 family, reading and generating text, images, video and audio in one model with a 1-million-token context for handling large codebases or long multi-session agent work.
Rather than training separate versions for coding, general agent tasks, visual understanding and cybersecurity, the creator trained one model across all of those tasks in a single mixed reinforcement-learning run, on the theory that skills learned in one area transfer and reinforce the others. It also uses an AI grader to compare multiple attempts at a task against each other rather than just marking each one pass or fail, to produce a more useful training signal.
You could use it to…
- Run one model trained across coding, agents, vision and audio
- Handle a million tokens of context for long agent sessions
- Use a checkpoint trained with an AI grader, not just pass/fail
Can I use this?
You'll need Multiple server GPUs and a serving framework · Setup needed
Worth knowing This is a large, resource-intensive model aimed at serious infrastructure rather than a single consumer GPU, and its reported gains from the mixed-training approach and the custom grading method are the creator's own technical-report claims. MIT licence. BuildTube has not run or verified this model.
Why it's here
Not trending this week
Last seen on the Hugging Face trending list on 28 September 2026. It stays on BuildTube so you can still find it.
- 658 people have liked it on Hugging Face.
- It was downloaded 89,231 times in the last 30 days.
Numbers checked on 6 October 2026; not refreshed since.
Behind it
See the code on Hugging FaceBuildTube has not run or verified this project. Everything above is written from what the creator published.
Made this?