GPT-Live - Full-duplex voice for ChatGPT

GPT-Live is OpenAI’s new full-duplex voice model for ChatGPT Voice. It can listen and speak at the same time, handle pauses and interruptions more naturally, and delegate harder search or reasoning work to frontier models in the background.

Add a comment

Replies

Best
Hi everyone! GPT-Live is OpenAI’s new voice model series now powering ChatGPT Voice. The important change is full-duplex interaction. It can listen and speak at the same time, so you don’t have to speak in perfect turns or wait for a clean pause every time. You can pause, interrupt, ask it to stay quiet, or let it give small listening signals without taking over the conversation. For harder work, GPT-Live can delegate to a frontier model in the background. At launch, that means GPT-5.5. Hopefully 5.6 soon :) So the voice layer keeps the conversation moving while search, reasoning, or more complex work happens behind the scenes. This feels closer to how a real voice assistant should work. One layer handles timing, listening, and flow. Another handles the deeper work when needed. Take out your phone and talk to ChatGPT Voice today - maybe you’ll get a small hint of what OAI’s future hardware interaction could feel like 😉

Natural interruptions and pause are something most voice assistants still struggle with. If GPT Live handles those well, it could make voice AI much more practical.

Finally a more natural sounding voice model.

Portuguese of Portugal still a bit choppy but perfectly usable!

the pause and interruption handling is what I've been most curious about — natural back-and-forth is still the hardest part of voice UX. building on the voice side too and the latency when delegating to a background reasoning model is the piece I'd love to stress-test. does it feel seamless in practice or do you notice a perceptible gap?

the listening-signal part is the interesting bit to me, not the interruption handling. most "full-duplex" demos look great one-on-one in a quiet room but the real test is a noisy kitchen or a car with the radio on, where the model has to decide what's actually speech directed at it vs background noise it should just ignore. curious if that discrimination holds up outside a controlled demo or if it still needs a wake word to reset context when it gets confused