LIVE
All stories ›
AI IN LIFENEWS
AI ModelsTools & AppsBusiness & DealsSocietyChips & ComputeResearchSafety & SecurityRegulation & PolicyRoboticsOpenAIAnthropicGoogle & DeepMindAlibaba / QwenxAIMetaByteDance
HomeOpenAI › TOOLS
TOOLS

GPT-Live-1 brings full-duplex voice to OpenAI's API

OpenAI says the model handles overlapping speech, follows instructions more reliably, and adds custom voices plus telephony for API builders.

GPT-Live-1 brings full-duplex voice to OpenAI's API
Symbolic illustration: a telephony gateway with blinking status lights next to a headset while an operator, seen from behind, adjusts the microphone arm.

In short

OpenAI has announced GPT-Live-1, an API speech model built for full-duplex conversation, stronger instruction following, custom voices, and telephony connections.

At a glance

  • GPT-Live-1 is announced as a speech model for the OpenAI API.
  • Headline capability: full-duplex conversation, with both sides able to speak at once.
  • Also listed: stronger instruction following, custom voices, telephony support.
  • The announcement page did not load for us; the server returned HTTP 403.
  • Pricing, model ID, languages, latency and availability are therefore unconfirmed.

OpenAI has introduced GPT-Live-1, a speech model it is putting into its API. The announcement names four capabilities: full-duplex conversation, stronger instruction following, custom voices, and telephony support.

Most voice assistants run half-duplex. They listen, detect a pause, then talk, which is why interrupting one usually means starting the turn over. Full-duplex means both sides can hold the floor at the same time, so a caller can cut in mid-sentence and the system is expected to respond to that interruption rather than finish its script.

On a support line, that is the difference between a conversation and a recorded announcement. How the model achieves it — barge-in detection, buffering, or end-to-end audio processing — is not something the material available to us describes.

Two of the four items are integration features rather than quality claims. Custom voices let a company keep one recognizable sound across its products instead of choosing from a fixed set. Telephony support points at wiring the model into the phone network, which is still where most real customer-service volume sits.

Which protocols, carriers or regions are covered is not stated in the summary we could read.

OpenAI's announcement page would not load for this report; the server answered with HTTP 403. What follows is therefore based on the announcement's title and summary line.

That leaves pricing, the exact API model identifier, supported languages, latency figures and the availability date unconfirmed. Only one independent newsroom has covered the release so far, so there was no second account to check the details against.

Voice in model APIs is not new territory. The recent bottleneck has been conversational behavior rather than audio quality: interruptions, clarifying questions, two people talking over each other. An announcement aimed squarely at that, and at reaching the phone network, reads as a pitch to call centers and voice-assistant builders.

The claim only becomes testable once the technical details are public. Until then, treat the feature list as described by OpenAI, not as verified here.

◈ AI-generated · sources linked

FAQ

What is GPT-Live-1?

A speech model OpenAI announced for its API, presented as a way to build more natural spoken conversations into applications.

What does full-duplex voice mean?

Both sides can speak at the same time, so the model is meant to handle being interrupted instead of waiting for a clean pause.

How much does GPT-Live-1 cost and when can developers use it?

We have no confirmed figures. The announcement page returned HTTP 403 when we tried to read it, so pricing and availability remain open.

Sources

More reports