Voice agents should become more than transcription tools. Instead of simply turning speech into text, they should understand what users are trying to achieve. The same spoken idea may need to become a short Slack message, a detailed email, a technical issue, or a personal note. A useful voice agent should adapt automatically.
This is the direction Loqua is exploring: voice as an intelligent layer across the apps people already use. It can understand the current application, nearby content, personal vocabulary, and writing style, then shape spoken thoughts into something ready to use. Users should be able to speak naturally without organizing every sentence beforehand.
The next step is real-time duplex conversation. A good agent should know when someone has finished speaking, when they are only pausing, and when they want to interrupt. It should also recognize signals beyond words, such as emotion, hesitation, emphasis, and tone. Its own voice should feel equally natural, expressive, and appropriate to the situation.
Behind this conversation, the agent needs deeper capabilities. It should remember relevant context, think through difficult requests, use tools, and continue working after the conversation moves on. The voice experience remains fast and fluid, while more complex reasoning happens quietly in the background.
We all know that long-horizon coding tasks are difficult, but understanding long, multi-turn conversations is even harder. The agent must remain consistent, respond with the right emotion, use tools correctly, and remember what matters.
Hey Product Hunt! 👋
I'm ARIA, the PM of Loqua.
A huge part of my work involves project coordination, so I spend most of my day jumping between apps — writing docs, reviewing charts, managing calendars, answering messages. And every single time, I'd think: why am I typing all of this? Why can't I just say what I want and have it done?
The models today can see your screen, understand context, reason about what you're looking at. But none of that power actually reaches you where you work. It's all trapped in chat windows.
That's what we're trying to change with Loqua. Bring the best of what frontier models can do — multimodal understanding, screen awareness, voice — and weave it directly into your workflow. So you can speak naturally, capture what's on your screen, ask questions about it, and go from thought to action without switching between tabs.
Our goal is simple: make working with computers feel as effortless as talking to a colleague. We're not there yet, but every update gets us closer.
Next step for Loqua is to continue integrating front abilities into the product while improving overall core response speed. This is already in progress.
I’ll be in the comments all day. Product collabs, feature expectations, feedback, or just general Q&A about existing features are all welcome.
I look forward to building an even stronger product together with you! Let’s talk! ❤️
The part about actually getting things done by voice sounds really interesting. What can it control right now?
Loqua
@trevor_loy23 Yep! that’s one of the areas where Loqua goes beyond voice typing. Right now on Mac, you can use voice to do things like create or edit reminders and calendar events, open apps or webpages, search Maps, take screenshots, and control things like Wi-Fi, Bluetooth, and AirDrop.
We’re gradually expanding what it can do, but the idea is simple: for small desktop tasks, you should be able to just say what you want instead of stopping what you’re doing to click through menus.
Loqua
@trevor_loy23 Theoretically it can control anything that linux commands do, we currently have calendar, notes, navigation, facetime messages, webpage search and so much more coming. Also doing a bunch of long horizon agentic control features, stay tuned for more!
This could be pretty handy for coding. Can I just talk through a change naturally and have it come out right?
Loqua
@reid_anderson4 Yep, that’s one of the use cases we’re excited about. You can talk through a change naturally, like “make this function cleaner” or “rewrite this to handle the error case,” and Loqua helps turn that into a more structured instruction or edit.
For very syntax-sensitive code, I’d still recommend reviewing the result, but for explaining changes, refactors, docs, or coding instructions, voice can be surprisingly fast.
Loqua
@reid_anderson4 Yes, we are working towards a more hands-free coding way! Try variable names, terms and complex logic, it's getting there!
Regarding the 'understand what's on your screen' feature, is that reading text, or real understanding of what application you're using?
Btw, Congratulations @joshua_zhou731 @ema_miao & team @Loqua 🚀✌️
@aymi_malik You can even download it and try it yourself — screenshot my profile pic and ask the AI to guess how old the kid in it is. 😄
Loqua
@ema_miao @aymi_malik Thanks buddy! By the nature of multimodal LLMs, it reads both text, visual and acoustic tokens. So it understands what application you are using, your snapshots and of course your voices, depends on how you use it. We are working on more agentic futures so it can get more long horizon tasks done.
@ema_miao @joshua_zhou731 That is an interesting methodology. Using multimodal tokens for such detailed contextual awareness is quite sensible. It is intriguing to learn that you are working on becoming more agentic and with a long-term horizon. Excited to see where that will lead. Wishing you continued success with the launch!
Loqua
@joshua_zhou731 @aymi_malik Really appreciate that 🙏 That long-term direction is probably what we’re most excited about too.
Loqua starts with voice typing and contextual understanding, but we’d love for it to gradually handle more complete workflows and longer-horizon tasks over time — without making the interaction feel heavier. Still early, but that’s definitely where we want to take it. Thanks again for the thoughtful comment and the support!
for app actions, can i see exactly what loqua heard and plans to do before it runs anything
Loqua
@mattkilmer Good question, you can see what Loqua understood, but we don’t currently show a full step-by-step action plan before every single action.
We’re putting a lot of thought into that boundary though, especially for actions with bigger consequences. The goal is for simple, low-risk things to stay fast, while anything more sensitive or ambiguous should be clear enough for you to confirm before Loqua acts.
@ema_miao that split makes sense, can users choose which actions always need confirmation or is the threshold fixed by loqua
Loqua
@mattkilmer That is a good idea for agent trust build, we will consider adding that feature for long horizon agentic functions, thanks!
If I’ve got a bunch of stuff open on screen, how does Loqua know what I’m actually talking about?
Loqua
@grant_w1 Good question. We don’t want Loqua to just guess based on everything that happens to be on your screen.
When there’s a lot going on, you can show/capture the relevant part and ask about that directly, which gives it much clearer context. Making screen context feel precise without adding more friction is something we’re still working on a lot.
workflow is actually kind of addictive.
Loqua
@jacob_hernandez4 Haha honestly that’s the best kind of feedback 😭 Once you get used to saying something and seeing it happen without breaking flow, going back to clicking around feels weirdly slow.
Loqua
@jacob_hernandez4 We are also planning workflow markets so everyone joins in, stay tuned!