Speech AI is a complete speech processing platform with three APIs: • Pronunciation Assessment — phoneme-level scoring, 17MB model, beats human experts • Speech-to-Text — word timestamps + confidence scores, same 17MB model • Text-to-Speech — 12 English voices via Kokoro-82M (#1 TTS Arena, Apache 2.0) All three ship as an MCP server with 8 tools, so AI agents can assess, transcribe, and speak in one integration. REST API and Azure Marketplace also available.
No reviews yetBe the first to leave a review for Speech AI Platform
Wispr Flow: Dictation That Works EverywhereStop typing. Start speaking. 4x faster.
Promoted
Maker
📌
Hi PH! We started with a pronunciation scoring API a few weeks ago and got great feedback. Since then we've expanded into a full speech AI platform.
The insight: language learning apps and AI tutoring agents need pronunciation scoring, speech-to-text, AND text-to-speech together. Having all three from one provider, with consistent latency and a single API key, is surprisingly rare.
The entire speech understanding stack (STT + pronunciation) shares a single 17MB model. TTS uses Kokoro-82M (115MB), which is currently #1 on the TTS Arena and Apache 2.0 licensed.
We also ship an MCP server with 8 tools so AI agents (Claude, GPT, etc.) can use all capabilities as tool calls — no custom integration needed.
Try the demo (3 tabs): https://huggingface.co/spaces/fa...
Would love your thoughts on what features to add next!