Run a 2B model built for phones that holds 131k tokens
Made by OpenBMB
View profile on Hugging Face (opens in a new tab)A small language model — 2.5 billion parameters in a standard Llama-style architecture — built specifically for running on a device rather than in a data centre, with a 131,072-token context that is unusual at this size.
The creator reports it reaching the best average in its size class on their comparison set and staying competitive with 4-billion-parameter models, with its clearest advantages in coding, mathematics, long-context work, tool use and agent tasks. Unusually, the team also published the training data behind it as open datasets, covering web text, tiered code data, agent examples and reinforcement learning samples.
Can I use this?
You'll need A phone, laptop or small GPU · Setup needed
Worth knowing Every benchmark figure is the creator’s own, measured against a comparison set they chose, and a small model’s strength on benchmarks does not always survive contact with an unusual real task. At this size it will be weaker on broad general knowledge than models several times larger, which is the trade being made. The repository holds several checkpoints from different training stages, so picking the wrong one gets you a model that has not been through its final tuning. Apache 2.0. BuildTube has not run or verified this model.
Why it's here
Not trending this week
Last seen on the Hugging Face trending list on 20 September 2026. It stays on BuildTube so you can still find it.
- 1,717 people have liked it on Hugging Face.
- It was downloaded 1,192,246 times in the last 30 days.
Numbers checked on 6 October 2026; not refreshed since.
Behind it
See the code on Hugging FaceBuildTube has not run or verified this project. Everything above is written from what the creator published.
Made this?