Voice agents should become more than transcription tools. Instead of simply turning speech into text, they should understand what users are trying to achieve. The same spoken idea may need to become a short Slack message, a detailed email, a technical issue, or a personal note. A useful voice agent should adapt automatically.
This is the direction Loqua is exploring: voice as an intelligent layer across the apps people already use. It can understand the current application, nearby content, personal vocabulary, and writing style, then shape spoken thoughts into something ready to use. Users should be able to speak naturally without organizing every sentence beforehand.
The next step is real-time duplex conversation. A good agent should know when someone has finished speaking, when they are only pausing, and when they want to interrupt. It should also recognize signals beyond words, such as emotion, hesitation, emphasis, and tone. Its own voice should feel equally natural, expressive, and appropriate to the situation.
Behind this conversation, the agent needs deeper capabilities. It should remember relevant context, think through difficult requests, use tools, and continue working after the conversation moves on. The voice experience remains fast and fluid, while more complex reasoning happens quietly in the background.
We all know that long-horizon coding tasks are difficult, but understanding long, multi-turn conversations is even harder. The agent must remain consistent, respond with the right emotion, use tools correctly, and remember what matters.
Loqua
Hi Product Hunt 👋 I'm Josh, founder of theloqua.ai.
I've spent my whole career as a speech and LLM researcher — and honestly, no voice product ever did what I actually needed. So I know exactly where ASR and voice agents break.
It got personal when carpal tunnel, TFCC, and an old ski injury turned long hours at the keyboard into real pain. I live at my desk, deep in conversation with LLMs all day, and typing simply couldn't keep up with how fast I think. I wanted to speak the way I think — not the way I type.
Not a wrapper
Old dictation was a cascade of small models — transcribing sounds but never understanding meaning, and one mis-heard word broke everything downstream. So our team — a group of research scientists who've spent years training and optimizing models — set out to build it properly. Loqua is not a wrapper. Instead of gluing everything to an LLM, we built one full omni model where your voice and your screen share the same context.
That single choice is what makes Loqua different. When you Capture to Ask — point at a chart, an error, or a table on your screen and just ask by voice — it reads code and tables more reliably, holds context better, and runs noticeably faster than bolt-on tools, thanks to a multimodal tokenizer and codec we designed ourselves. And because we train the model in-house, we keep sharpening it to real product needs and your feedback, every week.
Where we're headed
Voice and vision are only step one, though. The real direction is agentic — Loqua doing things for you. Say "remind me to send the Q4 draft at 3pm, and FaceTime Sarah" — and it happens, without you digging through menus or breaking your train of thought. Those little app-hunting detours are exactly what wreck your focus; Loqua takes them off your hands by voice so you can stay in the zone. And that's just our first attempt — next we're going after longer, multi-step tasks you can hand off in a blink and a simple sentence.
What you can do today
✍️ Speak, get polished text in any app
⚡ Capture to Ask — ask about anything on your screen
🎯 Get things done by voice across your Mac
🔊 Listen instead of read
🌐 ~100 languages, on macOS & Windows
Who it's for
It clicks fastest for people who live at a keyboard and now spend much of the day talking to AI — developers, PMs, researchers. A developer speaks a change straight to their coding agent in Cursor or Claude Code. A PM turns a rough spoken thought into a clean Slack update. Someone points at a chart on screen and just asks. And honestly, whatever you do all day — if your thinking outruns your hands, Loqua is for you.
AP, Yahoo Finance & PR Newswire just covered us: *"Voice-to-text was never enough. Meet voice that sees."*
https://apnews.com/press-release...
We also launched adn ranked 1st in Uneed last week.
https://www.uneed.best/weekly?da...
Your keyboard isn't going away — you'll just need it less.
🎁 Product Hunt Special — Stack your Pro days:
1. 30-day Pro trial — exclusive for PH users, find it in Promo code!
2. 14-day welcome trial — for every new user, automatically
3. +7 days each — share your code with a friend, you both get 7 extra days, no limit
We're especially trying to figure out which workflows people most want to make voice-native. So I'm curious — what's something that's annoying to type today but would feel natural to just say out loud? I'll be in the comments all day. 🙏
@joshua_zhou731 The Capture to Ask feature really caught my attention. Being able to point at something on your screen and ask about it by voice sounds really useful.
Loqua
@david_turner12 Hey David, thanks for coming! That is the real pain in our daily job working on tables and excels. The idea of making acoustic and visual context together for transformers brings the speed and accuracy another level, and we will keep improving it!
@joshua_zhou731 Yeah, I can see how useful that would be for tables and Excel. Being able to use your voice while pointing at what’s on screen makes things a lot easier. Looking forward to trying it out!
@joshua_zhou731 I like the idea of saying what’s on your mind instead of stopping to type it all out. That could save a lot of time during a busy workday.
Loqua
@oliver_graf1 Thanks Oliver! Most of this post is directly transcribed from pure oral input and the way of working has changed fundamentally. We will keep improving the model and user experience!
@joshua_zhou731 congrats on the launch, its a great product. it has many useful features, good luck with the lauch.
@joshua_zhou731 Oh, that’s actually a nice way to show the product 😄 Knowing the whole post was written through voice makes it even more interesting. I can see how this could make everyday work much easier. Excited to see what you build next!
@joshua_zhou731 nice launch! the screens and voice are context sounds really very useful,
what are the first multi step task you are the most excited to let Loqua handle by voice??
What made you build this instead of just using voice input with ChatGPT or another AI tool?
Loqua
@jack_pearson33 Great question. ChatGPT voice is great when you want to have a conversation with ChatGPT. What we wanted was voice to work inside the things you’re already doing — writing an email, editing text, looking at something on screen, coding, or handling a quick desktop task.
So Loqua is less about putting voice into another chat box, and more about making voice a layer across your desktop. That difference is really what pushed us to build it.
Loqua
@jack_pearson33 Very sharpe question! ChatGPT or ClaudeCode surely can be the overthrower for all apps. However we believe voice agent should be an universal core works anywhere on your computer/phone/wearable devices with a cheaper price, and that is how we want to make a difference. More features coming soon, stay tuned!
I'm curious where you have found voice is actually quicker than typing. Are there certain tasks where it's a clear win?
Loqua
@grace_gui02 Definitely not for everything. The clearest wins for us are when you already know what you want to say but don’t want to stop and type it all out: first drafts, long explanations, rewriting existing text, or capturing an idea before it disappears.
It’s also surprisingly useful when your hands are busy or when typing would force you to break focus and switch tools. That’s usually where voice starts to feel meaningfully faster, not just technically faster.
Loqua
@grace_gui02 We work with LLMs everyday and typing of long requirements is just so painful. Almost all tasks related to coding, PRD writing and docs editing won by voice dictation. A stats within our team shows 3-5x efficiency improvement in these tasks.
This really looks fantastic. I have been using voice to text apps for a bit which has immensely improved my workflow but they have never been able to read the screen to get context on what I was looking at. One question though, is it able to see the screen live or just with screenshots?
Look forward to giving this a try.
Loqua
@dmoniz22 Thanks, really appreciate that 🙏
Right now, Capture to Ask works from a screenshot/captured view rather than continuously watching your screen live.
We made that choice partly for precision and partly because we don’t want Loqua constantly observing everything on your screen in the background. You capture the context you want it to understand, then ask your question from there.
Would love to hear what you think once you try it!
Loqua
@dmoniz22 Currently it is triggered only through screenshots because we care about privacy, also the inference budget would be high if it is on all the time. We will introduce live chat, with live screen capture in the future, as our multimodal LLM keeps improving itself. Stay tuned!
Jinna.ai
Congrats on the launch! Building your own model sounds bold. Does it work locally on-device? Continuously? Wondering because I know at least on iPhones this would be really battery-draining
Loqua
@nikitaeverywhere Thanks Nikita! Our most capable models are server based so you don't have to worry about battery life. We are also working on more inference efficient on-device models so it suits variable network and using conditions.
how well does Loqua handle different acents and speech style when turning voice into polished writing?
Loqua
@johnnie_kuvalis Pretty well in our testing. Loqua is designed to handle different accents and more natural, conversational speech rather than expecting you to speak like you’re dictating word by word.
It also tries to clean up grammar and structure without flattening your original tone or speaking style. That said, very strong accents, noisy audio, or mixed-language speech can still be harder depending on the language, so we’re continuing to improve that. Curious which accent/language you’d want to use it with 👀
Loqua
@johnnie_kuvalis We used massive amount of training data and a multimodal LLM for this, so everything fits in one model, that is how we can manage accuracy, speed and different instruction requests better.
What has been the hardest part of making voice interactions reliable enough for real work?
Loqua
@albertnelson Honestly, the hardest part isn’t just speech recognition accuracy, it’s understanding intent reliably enough that the result is actually useful.
Real speech is messy: people pause, correct themselves, change direction mid-sentence, and rely on context that isn’t spoken out loud. For actions, there’s another layer too: knowing when Loqua should just do something versus when it should ask first. That trust boundary has been one of the most interesting problems for us.
Loqua
@albertnelson We worked pretty hard on some of the issues. Network latency for user around globe, multimodel & multilingual dictation accuracy, user query intention, and also device adaptation.