Give a model a million tokens of text and images at once
Made by DeepSeek
View profile on Hugging Face (opens in a new tab)A large model that reads images and text together and accepts contexts up to a million tokens — roughly a shelf of documents in one go.
Its design uses only a small slice of itself per word, which the creator says makes long, input-heavy work cheaper.
You could use it to…
- Hand it a whole contract bundle plus scanned appendices
- Ask which clauses changed since last year, in one pass
- Feed it a million tokens of text and images at once
Can I use this?
You'll need Server-class GPUs, or a hosted service · Setup needed
Worth knowing This is very large — the creator gives 552 billion parameters in the backbone — so running it yourself needs serious hardware and most people would reach it through a hosted provider. The efficiency claims are the creator’s own and come from its technical report. Released under the MIT licence. BuildTube has not run or verified this model.
Why it's here
Getting attention on Hugging Face
- 3,959 people have liked it on Hugging Face.
- It was downloaded 748,482 times in the last 30 days.
Numbers from the snapshot taken 1 October 2026; not refreshed since.
Behind it
See the code on Hugging FaceBuildTube has not run or verified this project. Everything above is written from what the creator published.
Made this?