Orchard - Speech-to-text for LatAm calls. Most accuracy per dollar.

by•
Speech-to-text API built for real Latin American Spanish calls. $0.025/hr batch, $0.036/hr with word-level speaker diarization. OpenAI-compatible: swap one line. We benchmarked 7 engines on real calls and published audio, references and code.

Add a comment

Replies

Best
Hey Product Hunt 👋 I'm Mateo, founder of Orchard. Every speech-to-text benchmark we found measured English, read in a studio, one speaker, good mic. Our users transcribe call center calls in Colombia, Argentina, Mexico and Peru: two people talking over each other, IVR in the background, noisy lines. Those rankings predicted nothing. So we built our own: 13 real Latin American calls, 7 engines, audio, references and scoring code all public. We came 4th overall. 3 WER points from #1. We published it anyway, because a benchmark where the publisher wins everything helps nobody. But the overall number hides the part we care about most. On the 4 held-out calls, the hardest set: • Hand-transcribed and never published, so no engine could have trained on them • Two people, turns overlapping, up to 62 turns in 5 minutes • Real recorded calls from Argentina and Colombia On those, Orchard is #2 of 7, and #1 on the longest, messiest one (5 minutes, 62 turns, 1,012 words). That's the audio our users actually send. What we optimized for is accuracy per dollar: • $0.025/hr for batch transcription • $0.036/hr with word-level speaker diarization included • 2.7x more accuracy per dollar than Soniox, 4.6x AssemblyAI, 17x Deepgram (list prices, diarization included) It's OpenAI-compatible, so switching is one line (baseURL). There's also a TypeScript SDK, a Vercel AI SDK provider, an MCP server and n8n templates. We never train on your audio. What it's not (yet): no real-time streaming, we're batch only. Speaker attribution on those hard calls is where we're still behind the leaders, and it's the next thing we're fixing. And our Portuguese is behind our Spanish. 500 free minutes on signup, no card. What I'd love from you today: throw your hardest audio at it (bad line, crosstalk, heavy accent) and tell me where it breaks. I'll be here all day answering everything.