SpeechifyAI Simba Voice Agents - Voice agents powered by Simba 3.2 the world's #1 voice model

Simba Voice Agents: build production voice agents on Simba 3.2, the #1 model on Artificial Analysis. Sub-100ms, streaming-native, real emotion and SSML. The best real-time voice, now a full agent platform, on Speechify's new Speechify Developer Platform.

Add a comment

Replies

Best

Hey folks! Luke here, I run developer relations at Speechify.


This is the underdog story for API providers. Speechify has spent years making our models run efficiently because our consumer business demanded it, tens of millions of listeners with some of the best voices on the planet. Commercially, the faster and more efficient it got, the better. And that work is why we can now put the best-rated model in the world on our API.


Simba 3.2 just went #1 on Artificial Analysis, and on Voice Arena it's the top real-time model for both quality and price, at $6 per 1M characters, the cheapest in the top ten. Capable of sub-100ms latency, streaming-native, with real emotion and SSML control. It's the same model that powers our consumer apps for 60M+ people, now rolling out on our new developer platform across SpeechifyAI Agents and SpeechifyAI Build, with a REST API and first-party TypeScript and Python SDKs.


Most labs built for the benchmark and priced for the enterprise. We built for listeners and priced for production.

 Congrats on the launch! I liked the story behind how this product evolved.

With new AI products launching almost every day, I'm curious how you're thinking about marketing. What's your strategy to make Speechify stand out when many products are making similar claims?

 Thank you! I understand what you're saying, but right now they're not just claims, we're literally ranking top on Artificial Analysis which is notoriously difficult to do. It's a blind-tested by community members.

But, it is going to be a battle. Right now we're focussed on reliability and visibility. We want everyone to know we're here, and when they give us a go it has to work.

Our quality is up there at the top and our price is the lowest in the top ten, so if folks are looking for quality we want to prove it is no-longer a decision you make with your budget, but with your ears

Speechify being framed as an AI Voice Assistant makes me wonder about the core workflow you’re optimizing for. Is the main use case more around dictation, hands-free task execution, or connecting voice commands into AI workflow automation? Since it’s also listed near Developer Tools and AI Agents, I’d be interested to know whether there are integrations or APIs planned for teams that want to plug it into existing tools.

 Speechify has for the last almost-decade been focussed on providing really great consumer apps. This launch is part of a larger step into APIs and integrations. You can checkout more on our developer site

Congrats to the Speechify team! 🎉 Loved seeing this expand into a full developer platform.

Curious about voice customization on : can developers bring their own custom-cloned voices or fine-tuned emotion profiles onto the Simba 3.2 API, or is it currently optimized for Speechify's core library of voice models?

 we support 3.2 cloning but it is through our FDEs currently while we ensure you're getting the highest quality clones. 3.0 and 1.6 cloning are self-serve. You can jump straight into a conversation about that with our team at

Anywhere I can hear a sample of this voice?

 Sure can, the hero on is blind testing, you'll hear us and a random competitor

 I feel its heavy leaning towards audio-book and podcast. I guess because it was trained on that? Which serves 2 usecases. But voice has so many other usecases. like product presentation videos etc. It would be interesting to hear some examples for that too.

 Speechify started in that space, but is a Voice API provider with no intended vertical or specific usecase. Artificial Analysis is a leaderboard that favours no particular usecase either, and blind tests on user provided strings instead of any pre-defined strings - impossible to game.

 Right but all the examples on the landingpage are podcast or audiobook. 😅 Maybe add other examples for different usecases?

Congratulations on the launch🎉🎉🎉. Really liked the comparison table, and absolutely loved the emotion control displayed on the website, can we also get different dialects of different languages like in arabic there are many dialects, so are there any options for that ?

 we're rapidly working on more languages for Simba 3.0 multi-lingual! If you'd like to discuss Arabic and our language pipeline then I'd recommend jumping in a call with our team

 👍👍👍

Congrats on the launch. Do you have any plans to make the speech-to-text/dictation feature available through APIs in the future? I'd love to be able to transcribe user speech directly inside my own app

 I'm going to double check what I am allowed to say here ;)

The way Speechify slides into Google Docs and Gmail without breaking your flow is genuinely clever, the keyboard shortcut overlay feels thoughtful and not at all clunky.

 Our consumer app is beautifully polished, I've used it for way longer than I've been here and I love it. I was so excited to join Speechify's move into APIs, and I hope folks get to build their own amazing experiences with our high quality voices!

Fantastic! I'm checking it out right now. I'm currently using one of your competitors, but a quick glance at your site proves that a closer look is warranted. I need TTS and voice-agent. Congratulations on your launch!

 hey Terry, thank you so much! I'd love to see you get in touch with our team to see if we can help. Our agents are now powered by Simba 3.2, we have TTS and cloning if you're looking for a bespoke voice at our level of quality

Congrats to the entire team on the launch. Excited for people to start seeing Speechify as a developer & B2B platform beyond our consumer roots. We've learned so much about building a cost efficient, high quality, low latency voice model and serving it globally from our time in consumer and now pumped for other businesses and developers to make use of 5+ years of AI research we've been driving in-house.

sub-100ms at $6/1M chars is a serious combo. we're chasing similar latency for voice checkins but staying fully on-device instead of hitting an api — is that number end-to-end from audio in to first byte out, or just model inference?

 we have a paper on sub-100ms that explains it all, but it would be part of a conversation with our team. You can book time with someone here:

12
Next