BuildTube Preview About
Back to the feed

Give a model a million tokens of text and images at once

A large model that reads images and text together and accepts contexts up to a million tokens — roughly a shelf of documents in one go.

Its design uses only a small slice of itself per word, which the creator says makes long, input-heavy work cheaper.

You could use it to…

  • Hand it a whole contract bundle plus scanned appendices
  • Ask which clauses changed since last year, in one pass
  • Feed it a million tokens of text and images at once

Can I use this?

Needs server-class GPUs

You'll need Server-class GPUs, or a hosted service · Setup needed

Worth knowing This is very large — the creator gives 552 billion parameters in the backbone — so running it yourself needs serious hardware and most people would reach it through a hosted provider. The efficiency claims are the creator’s own and come from its technical report. Released under the MIT licence. BuildTube has not run or verified this model.

Why it's here

Getting attention on Hugging Face

  • 3,959 people have liked it on Hugging Face.
  • It was downloaded 748,482 times in the last 30 days.

Numbers from the snapshot taken 1 October 2026; not refreshed since.

Behind it

See the code on Hugging Face

BuildTube has not run or verified this project. Everything above is written from what the creator published.

Made this?

Back to the feed