BuildTube Preview About
Back to the feed

Turn document scans and photos into structured text

TeleOCR is a lightweight vision-language model — about 1.2 billion parameters (the model's learned settings) — built for reading documents: give it scans or camera photos of pages and it returns their text, tables, formulas, and reading order.

It handles clean digital files and messy real-world captures alike, like tilted phone photos and warped pages, using geometry-aware modeling that understands page curvature. It runs in Python through the transformers library; the team also publishes a vLLM variant, and there is a community GGUF conversion for llama.cpp. The creators report it topped the ICDAR 2026 Sci-ImageMiner challenge and scored 96.87 overall on the OmniDocBench document benchmark, ahead of the rival models they compared against.

Can I use this?

Needs a model runtime on your computer

You'll need Python 3.10+, the transformers library, and a CUDA-enabled GPU · Setup needed

Worth knowing The card describes TeleOCR as open-source but names no specific licence, so check the repository before commercial use; like any document reader it can misread unusual layouts, and practical inference needs a CUDA-enabled GPU. BuildTube has not run or verified this model.

Why it's here

  • 1,143 people have liked it on Hugging Face.
  • It was downloaded 31,584 times in the last 30 days.

Numbers from the snapshot taken 1 October 2026; not refreshed since.

Behind it

See the code on Hugging Face

BuildTube has not run or verified this project. Everything above is written from what the creator published.

Made this?

Back to the feed