Dictation API turns a spoken clip into finished text. Filler and false starts come out, and the output takes whatever shape you instruct: notes, a commit message, a reply to a customer. Built on Universal-3.5 Pro. 19 languages, under a second on short clips, $0.62/hr all in.
Introducing Universal-2: The latest advancement in Speech-to-Text technology. Capture the complexity of human speech, enhanced transcript quality, and better conversational insights by tapping into the next generation of Speech AI.
Universal-3.5 Pro is AssemblyAI's most accurate speech-to-text model, now available at our Realtime & Async endpoints. It transcribes every conversation exactly as it's heard—code-switching across 18 languages, our most accurate speaker diarization yet, and contextual prompting to steer results.
Universal-3 Pro is a new class of speech language model built for Voice AI. Control transcription using instructions and domain context like names, terminology, and topics to get accurate output at the source. No custom models, no post-processing pipelines, no hallucinations. Includes 1,000 keyterms, audio tagging, and 6-language code-switching for $0.21/hr.
Try AssemblyAI's most capable and highly trained speech recognition model trained on 12.5M hours of multilingual audio data. Universal-1 achieves best-in-class speech-to-text accuracy, reduces word error rate and hallucinations, and improves timestamps.
The fastest path to a working Voice Agent, built on the most accurate Voice AI in the market. Stream audio in, get audio back. We handle the rest.
~1s latency. Best-in-class accuracy on the stuff that matters (numbers, emails, names). Tool calling that doesn't go silent. Mid-call prompt + voice + tool updates. $4.50/hr flat. No per-token. No concurrency caps.
Most devs ship a working agent the same day.
Universal-Streaming delivers all the streaming speech-to-text voice agents need in one robust API: ultra-fast immutable transcripts, higher accuracy, built-in endpointing, and transparent pricing at $0.15/hour with unlimited concurrency.
Universal-3 Pro Streaming is the most accurate real-time STT model for voice agents. With entity detection, speaker labels, and code switching, it's built for the hard stuff: disfluencies, alphanumerics, and noisy environments. One API. 99+ languages. Try it free.