Voice & AI audio
OpenAI's new voice AI, explained — what it can do now
Smarter conversations, live translation and instant transcription. Here's the plain-English version.
The answer
On 7 May 2026 OpenAI launched voice AI that reasons, translates and transcribes live.
If you've ever tried to have a real conversation with a voice assistant and ended up repeating yourself, simplifying your question, or just giving up — this update is part of why that's about to change. On 7 May 2026, OpenAI added three new voice tools to its developer platform, and they work together to make AI a genuinely smarter, more useful thing to talk to. Here's what's changed, what each tool does, and where you'll actually notice it.
The three new tools, explained simply
GPT-Realtime-2 — the one that actually thinks. This is the most important of the three. Previous voice AI could understand what you said but struggled the moment you asked for something that required real reasoning — checking multiple things, holding context across a conversation, working through a problem step by step. GPT-Realtime-2 brings GPT-5-class reasoning into a live voice model, which means it can handle harder requests mid-conversation without losing the thread. Think of the difference between a scripted phone menu and an assistant who genuinely listens and thinks.
GPT-Realtime-Translate — live speech translation. This one translates speech from 70+ input languages into 13 output languages in near real time — while you're still talking. That makes it useful for travel, international calls, multilingual customer service, and any situation where two people speak different languages and need to understand each other without delay. It's priced per minute, which is how services like this usually work.
GPT-Realtime-Whisper — instant transcription. This streams what you say into written text as you speak — no wait, no batch upload, just live text as the words come out. It's built on OpenAI's well-regarded Whisper transcription technology, now available as a live streaming service. Useful for meeting notes, accessibility tools, live captions, and any app that needs spoken words turned into text fast.
OpenAI advanced voice intelligence on 7 May 2026 with three new Realtime API models: GPT-Realtime-2 for reasoning-capable voice, GPT-Realtime-Translate for live multilingual translation, and GPT-Realtime-Whisper for streaming transcription.
Where you'll notice this
You won't download these tools yourself — they're for companies and developers who build apps. But you'll feel the difference in the products you use. The smarter voice assistants arriving in customer service apps this year are likely running on something like this. Meeting apps that transcribe as people speak, travel apps that translate conversations live, AI assistants that can actually handle a multi-step request without getting confused — these are all downstream of what OpenAI shipped on 7 May.
Here's a quick look at the three models and who they're most useful for:
| Tool | What it does | Best for |
|---|---|---|
| GPT-Realtime-2 | Smarter spoken conversations | AI assistants, customer support, complex voice apps |
| GPT-Realtime-Translate | Live speech translation, 70+ languages in → 13 out | Travel, international calls, multilingual service |
| GPT-Realtime-Whisper | Real-time speech-to-text | Meetings, accessibility, live captions |
Think of it as the plumbing behind the better voice experiences arriving across lots of apps this year.
What to be honest about
GPT-Realtime-2 is billed by token — unlike Translate and Whisper which are billed per minute — reflecting the variable computational cost of real-time reasoning in voice conversations.
The bottom line: talking to AI is getting meaningfully better, and OpenAI's May update is a big reason why. You don't need to do anything — the improvement will come through the apps you already use, as developers build on these new tools. If you've been waiting for voice AI that actually understands a proper conversation, that wait is getting shorter.
Frequently asked questions
Can I use these voice features now?
Is the live translation any good?
How is GPT-Realtime-2 different from the voice AI I've used before?
Will this make voice assistants like Siri or Alexa better?
Is GPT-Realtime-Whisper the same as the Whisper I can already use?
Sources
- Advancing voice intelligence with new models in the API — OpenAI, 7 May 2026
- Realtime API guide — voice agents, translation, transcription and speech models — OpenAI Platform Docs, 7 May 2026
- OpenAI launches new voice intelligence features in its API — TechCrunch, 7 May 2026