Turn document scans and photos into structured text
Made by XingChen-AGI
View profile on Hugging Face (opens in a new tab)TeleOCR is a lightweight vision-language model — about 1.2 billion parameters (the model's learned settings) — built for reading documents: give it scans or camera photos of pages and it returns their text, tables, formulas, and reading order.
It handles clean digital files and messy real-world captures alike, like tilted phone photos and warped pages, using geometry-aware modeling that understands page curvature. It runs in Python through the transformers library; the team also publishes a vLLM variant, and there is a community GGUF conversion for llama.cpp. The creators report it topped the ICDAR 2026 Sci-ImageMiner challenge and scored 96.87 overall on the OmniDocBench document benchmark, ahead of the rival models they compared against.
Can I use this?
You'll need Python 3.10+, the transformers library, and a CUDA-enabled GPU · Setup needed
Worth knowing The card describes TeleOCR as open-source but names no specific licence, so check the repository before commercial use; like any document reader it can misread unusual layouts, and practical inference needs a CUDA-enabled GPU. BuildTube has not run or verified this model.
Why it's here
- 1,143 people have liked it on Hugging Face.
- It was downloaded 31,584 times in the last 30 days.
Numbers from the snapshot taken 1 October 2026; not refreshed since.
Behind it
See the code on Hugging FaceBuildTube has not run or verified this project. Everything above is written from what the creator published.
Made this?