Dictation API turns a spoken clip into finished text. Filler and false starts come out, and the output takes whatever shape you instruct: notes, a commit message, a reply to a customer. Built on Universal-3.5 Pro. 19 languages, under a second on short clips, $0.62/hr all in.
Universal-3.5 Pro is AssemblyAI's most accurate speech-to-text model, now available at our Realtime & Async endpoints. It transcribes every conversation exactly as it's heard—code-switching across 18 languages, our most accurate speaker diarization yet, and contextual prompting to steer results.
Universal-Streaming delivers all the streaming speech-to-text voice agents need in one robust API: ultra-fast immutable transcripts, higher accuracy, built-in endpointing, and transparent pricing at $0.15/hour with unlimited concurrency.