BuildTube Preview About
Back to the feed

Picture published by the creator on Hugging Face.

Transcribe speech live, continuously, with no length limit

Audio8 ASR Infinite is a speech-to-text model built for real-time, unlimited-length transcription rather than processing a finished audio file after the fact.

Most streaming transcription models drift out of sync or need to be restarted for very long audio; this one uses a "rolling" memory technique that keeps both memory use and response delay constant even running 24 hours a day. It supports Chinese and English, lets you choose how often it emits new text (roughly 6 to 12.5 times per second) and how much delay to trade for accuracy, and includes logic to tell a thinking pause or a stutter apart from someone actually finishing what they were saying.

You could use it to…

  • Transcribe speech live with no length limit, 24/7
  • Build a captioning feature that never drifts out of sync
  • Tell a thinking pause apart from someone finishing a sentence

Can I use this?

Needs a GPU

You'll need A GPU and the transformers library, or vLLM for 24/7 use · Setup needed

Worth knowing This is a preview release described by the creator as delivering only the transcription base, with a further "realtime semantic perception" capability still being built; accuracy figures in the model card are the creator's own reported error rates. Apache-2.0 licence. BuildTube has not run or verified this model.

Why it's here

Trending on Hugging Face

  • 2,023 people have liked it on Hugging Face.
  • It was downloaded 27,228 times in the last 30 days.

Numbers from the snapshot taken 1 October 2026; not refreshed since.

Behind it

See the code on Hugging Face

BuildTube has not run or verified this project. Everything above is written from what the creator published.

Made this?

Back to the feed