GPT-Live: OpenAI's Real-Time Voice AI, Explained Simply
OpenAI's GPT-Live now powers ChatGPT Voice, and it works fundamentally differently from the voice assistants before it: it can listen and speak at the same time, responding in well under 300 milliseconds. Here's what that means for anyone talking to ChatGPT.
โก Quick facts
- What it is: A full-duplex voice AI system, replacing ChatGPT's older Advanced Voice Mode
- Speed: End-to-end responses under 300 milliseconds
- Two versions: GPT-Live-1 (Go/Plus/Pro users) and GPT-Live-1 mini (Free users)
- Available on: ChatGPT Voice across iOS, Android, and the web
What "full-duplex" actually means
Older voice assistants, including ChatGPT's previous Advanced Voice Mode, worked in turns: you spoke, it processed your full sentence, then it replied โ a cycle closer to a walkie-talkie than a conversation. GPT-Live instead runs on a full-duplex architecture, processing what you're saying and generating its own response simultaneously. Many times per second, it decides whether to keep listening, jump in, pause, or quietly acknowledge you with something like "mhmm" while you're still talking โ much closer to how two people actually talk to each other.
Why sub-300ms latency matters
Human conversation has natural gaps of well under a second between one person finishing and the other responding; anything slower starts to feel like a phone call with a lag. By cutting end-to-end response time to under 300 milliseconds, OpenAI is aiming to close that gap so a voice conversation with ChatGPT feels closer to talking with a person than issuing a command and waiting.
Two versions, two audiences
OpenAI shipped two models: GPT-Live-1, which became the default for ChatGPT Go, Plus, and Pro subscribers, and a lighter GPT-Live-1 mini, which became the default for Free users, replacing Advanced Voice Mode entirely for that tier. The same push toward everyday agents is why OpenAI agents elsewhere keep making headlines.
How it handles hard questions
Keeping a conversation fast doesn't leave much time for deep reasoning mid-sentence, so GPT-Live uses a two-layer design: a fast interaction layer manages the live back-and-forth, while a separate delegation layer quietly hands genuinely complex questions off to GPT-5.5 in the background โ so you keep talking naturally while the harder thinking happens out of sight.
Why this matters
Voice has long been the most natural way people communicate, but AI voice assistants have mostly felt stilted and turn-based. A model built specifically to listen and speak at once, at near-human latency, points toward voice becoming a genuinely comfortable way to use AI day-to-day โ for quick questions, hands-free tasks, or just thinking out loud with an assistant that doesn't wait for you to finish before it starts helping.
Frequently asked questions
What is GPT-Live?
OpenAI's full-duplex voice AI system now powering ChatGPT Voice -- it can listen and speak at the same time rather than waiting for turns.
How fast does GPT-Live respond?
End-to-end response times are under 300 milliseconds, close to natural human conversation pace.
What's the difference between GPT-Live-1 and GPT-Live-1 mini?
GPT-Live-1 is the default for Go, Plus, and Pro users; GPT-Live-1 mini is the default for Free users, replacing Advanced Voice Mode.
How does GPT-Live handle complex questions?
A fast interaction layer keeps the conversation flowing while a delegation layer hands harder reasoning to GPT-5.5 in the background.
Where is GPT-Live available?
It powers ChatGPT Voice on iOS, Android, and the web.