Lifestyle

GPT-Live 1 Comes to the API: Can It Take Customer Service Calls, Whose Voice Is It, Will It Record?

On September 10, 2026, OpenAI’s API changelog noted that voice model GPT-Live 1 is now generally available in the API, so developers can wire it into phone support, apps, or their own products. Drawing on OpenAI’s developer documentation, this article covers the model’s API specs, how calls connect, the documentation’s note that outbound calls are not supported, its 12 named voices and unlisted spoken languages, and how recording and per-minute billing work; this site has not tested it.

About 13 min read

Original illustration: a phone with sound-wave lines; one line bends toward a rounded panel with a speech bubble and sound waves, showing an incoming call answered by a voice model.
Image: Mokaair (© Mokaair)

On September 10, 2026, OpenAI’s API developer documentation logged a new changelog entry: the full-duplex voice model GPT-Live 1 is now generally available in the API, under the model ID gpt-live-1 and the endpoint v1/live/sessions. The official documentation states that voice sessions cost $0.05 per minute, billed per second, with backend model and tool usage charged separately. What this means is that a voice model that can listen and speak at the same time is a component developers can buy and build into their own phone systems, apps, or products.

This article was checked on September 18, 2026 against OpenAI’s developer documentation: the API changelog, the GPT-Live 1 model page, the telephony and SIP integration guide, and the session management guide, reflecting the versions available on the day of checking; these documentation pages carry no version numbers and their content may change later. This site did not actually integrate with or test this API; the capability descriptions in this article are OpenAI’s own statements. This article covers only this update on the API developer side, not the plans or settings on the ChatGPT consumer side.

The September 10 Changelog Entry: Generally Available, the Model ID, and Two Delegation Modes

The official model page defines GPT-Live 1 as a full-duplex voice model: it can listen and speak at the same time, delegating reasoning and tool use to a backend agent — meaning that when something needs to be looked up or computed, that work is handed off in the background while the conversation continues. The changelog describes the September 10 status as “now generally available in the API”; that phrase describes an availability status, and the four documents this article checked do not say this is the first time it has appeared. The telephony integration guide, in fact, notes that existing integrations may still receive a deprecated incoming-call event name, which suggests integrations were already running before this date.

The model page lists the model ID as gpt-live-1, with the corresponding endpoint v1/live/sessions; the official documentation describes it as “our premier model for natural, expressive voice conversations with smooth interruption handling” — this is OpenAI’s own characterization. Both input and output modalities are audio and text; image and video are not supported. The model page lists a knowledge cutoff of July 31, 2025, which is the point in time for the model’s own training data.

When developers want GPT-Live 1 to handle questions that require looking things up or computation, the documentation describes two delegation modes: Responses delegation, paired with an OpenAI model, and client delegation, which connects to a backend the developer runs themselves; if delegation is omitted or set to null, it defaults to client mode. The documentation states that the delegation mode cannot be changed once a session has started — changing it requires starting a new session.

The Model’s API Specifications: Only One Supported Endpoint, No Free Tier

GPT-Live 1 has only one endpoint marked as supported in the API — Live (v1/live/sessions); the official endpoint list marks existing endpoints such as Chat Completions, Responses, and Realtime, along with the older Completions (legacy), as not supported. The telephony integration guide also notes that each API has its own authentication, session creation, and event contract, and they cannot be substituted for one another.

Usage is also measured differently: the documentation states that rate limits are measured in concurrent sessions, and lists Free as an unsupported usage tier. Depending on the account’s usage tier, Tier 1 through Tier 5 can open 25, 50, 200, 300, and 500 concurrent sessions respectively.

Compiled from OpenAI developer documentation; checked on September 18, 2026.
ItemDetailsNotes
Model IDgpt-live-1Only supported endpoint: v1/live/sessions
Supported modalitiesAudio and text (both input and output)Image and video not supported
Delegation modeResponses delegation or client delegationSelected when the session is created; cannot be changed mid-session
Concurrency limitTier 1–5: 25 / 50 / 200 / 300 / 500Free tier not supported
Voice pricing$0.05 per minute, billed per secondBackend model and tools billed separately

When a Call Comes In: How GPT-Live 1 Takes Over

To connect GPT-Live 1 to a phone system, the official documentation lists two paths: Direct SIP, where the telephony provider exchanges call audio with OpenAI while the developer’s application handles webhooks, session configuration, call decisions, and business logic; and a server audio bridge, where the developer’s application relays audio from the telephony provider or a meeting room to GPT-Live over WebSocket, managing both connections, event translation, playback, and the call lifecycle itself. The documentation mentions that Twilio, Telnyx, LiveKit, and Daily/Pipecat each have dedicated integration guides.

Once a call comes in: the developer’s application applies its own authorization and routing rules, using a separate API action to accept or reject the call — accepting requires sending settings such as the model, voice, and delegation mode, while rejecting requires attaching a SIP status code between 300 and 699 (486, for busy, for example); the first accept or reject decision wins, and any decision sent afterward is rejected. Two further actions are available during a call — transferring it elsewhere or hanging up directly — both of which return 200 OK with no content on success.

At the end of this section, the documentation states plainly that this flow handles inbound calls, and does not support creating outbound SIP calls through POST /v1/live/sessions; for outbound calling, the documentation directs developers to use a partner integration instead, describing outbound calling as something the provider owns. The documentation also warns that SIP headers attached to an incoming call should be treated as untrusted caller metadata, not as a basis for authorization.

Four-panel diagram: four key points about GPT-Live 1 in the API — calls are inbound-only, supported languages are not listed, recording is off by default, voice costs $0.05 per minute
Four key points about GPT-Live 1 in the API: calls are inbound-only, voice languages are not listed, recording is off by default, and pricing is $0.05 per minute; compiled from OpenAI developer documentation, checked on September 18, 2026. · Image: Mokaair (© Mokaair)

Whose Voice Is It: 12 Built-in Options, but These Four Documents Don’t List Supported Languages

GPT-Live 1’s voice isn’t limited to just one option. Besides the default, marin, the official documentation lists 12 additional named voices to choose from — for example, vesper, whose language field is marked English and regional-influence field British, and gleam, marked English and North American; as well as bossa and tempo, both marked Portuguese with a Brazilian regional influence. Developers select one and put it in the session settings when creating a session; once a conversation has started, it cannot be changed mid-call — changing it requires starting a new session.

What readers in Taiwan should note: the documentation uses “regional influence” to describe these voices, while specifically stating that this refers only to speaking style and is not a guarantee of accent fidelity. More importantly, the four official documents this article checked do not list which spoken languages this model supports — the language field for the 12 named voices shows only English and Portuguese, and the default voice, marin, has no language specified either; when the documentation discusses greetings, it tells developers to use the language specified by the application until the caller speaks, and to test with the languages their own application supports. In other words, if GPT-Live 1 really is answering a customer service call, how well it handles Chinese is a question these four documents do not answer — it should not be assumed that it can.

The official documentation also mentions that, once approved, a custom voice can be built from your own recordings; the details are in a separate guide that this article does not go into.

Is It Recorded, How Long Does It Remember, and What Does a Few Minutes on the Phone Cost

Whether a call is recorded is up to the developer — it is not on by default. The official documentation states that a session’s recording setting defaults to off; the developer must turn it on, and the project must first be approved, before a completed recording can be kept for download or “forking” into a new session, and downloading or forking also requires a data policy that permits persistence. Saved recordings expire after 30 days; if a project has Zero Data Retention enabled, the recording setting is always treated as off, making downloads and forks unavailable. A downloaded recording is a stereo WAV file, with the input and output audio in the left and right channels respectively.

There’s also a cap on memory for long calls. The documentation states that a session’s default context window holds 128,000 tokens, which includes the instructions given by the developer, the conversation text, and audio tokens that do not appear in the transcript; when usage exceeds 90% of that window, the system starts a replacement voice engine within the same session, and the new engine receives the original instructions plus up to 8,192 tokens of conversation history (recent messages, and, when available, a summary of older messages). That means that in a very long call, details discussed earlier may be summarized or omitted.

The pricing formula itself is simple: voice sessions cost $0.05 per minute, billed by actual seconds, and the documentation notes that this is not rounded up to the next whole minute; the backend model and tools are billed separately at their own rates, and these four documents do not combine the two into a single worked example. What follows is this article’s own conversion at the rate checked on September 18, 2026, not an official figure: for three full minutes of voice alone, that works out to $0.15; the actual total also depends on how much work the backend did. When a session ends, the official documentation lists reasons including the application closing it, the session reaching its duration limit, a safety filter stopping it, the remote connection ending normally, and the connection being interrupted unexpectedly — but the documentation does not state what the duration limit actually is in minutes.

Frequently asked questions

Was September 10, 2026 the first time GPT-Live 1 became usable?

The changelog entry says it “is now generally available in the API”; that phrase describes an availability status, and the four official documents this article checked do not say this is the first time it has appeared. The telephony and SIP integration guide, in fact, mentions that existing integrations may still receive a deprecated incoming-call event name, which indicates integrations were already running before this date. These four documents also do not describe an earlier preview period or a staged rollout schedule.

Does GPT-Live 1 understand Chinese? Can you call in and ask questions in Chinese?

The four official documents this article checked do not list which spoken languages this model supports. Of the 12 named voices listed in the documentation, the language field shows only English and Portuguese, and the default voice, marin, has no language specified either; when the documentation discusses greetings, it tells developers to use the language specified by the application until the caller speaks, and to test with the languages their own application supports. As of the documentation version this article checked on September 18, 2026, there is no statement confirming Chinese support, so it should not be assumed that it can understand or respond fluently in Chinese.

Can developers in Taiwan use this API right now?

The four official documents this article checked do not list which countries or regions can use it, and none of them mention Taiwan specifically. Whether you can apply for or use it should be based on what your own account actually shows on the OpenAI platform; this article has not verified whether Taiwan-based accounts can currently use this feature.

If a business wants to connect this to its phone system, is OpenAI’s own channel the only option?

Not necessarily. The official documentation states that this call flow handles inbound calls, and does not support creating outbound SIP calls through POST /v1/live/sessions; for outbound calling, the documentation directs developers to use a partner integration instead, naming Twilio, Telnyx, LiveKit, and Daily/Pipecat. These fall under each third-party provider’s own service scope, so actual features and pricing depend on what each provider states.

Is this voice session recorded and retained?

Not by default. The official documentation states that a session’s recording setting defaults to off; the developer must turn it on, the project must first be approved, and downloading or forking also requires a data policy that permits persistence before any recording is kept; saved recordings expire after 30 days. If a project has Zero Data Retention enabled, the recording setting is always treated as off, and there is no recording available to download.

Roughly how much does a call cost?

What’s officially published is the rate for voice itself: $0.05 per minute, billed by actual seconds, as checked on September 18, 2026. What follows is this article’s own conversion, not an official figure: a three-minute call, counting voice alone, works out to $0.15. The backend model and any tools it calls are billed separately at their own rates, and the actual total for a call depends on how much work it did; these four documents do not provide a total-cost example for an entire call.

  • Lifestyle

    Gemini 3.8 Live and Extended Thinking: New Voice Models, Can Your Account Use Them?

    On September 15, 2026, Google introduced two voice-dialogue models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Drawing on Google's official model post, developer post, model card, and pricing page, this article lays out what developers, enterprises, everyday users, and Workspace subscribers can each access, how the pricing works, the knowledge cutoff date the model card states, and the availability regions the official pages do not specify.

  • Lifestyle

    GPT-Live Lets ChatGPT Voice Listen as It Talks: Interruptions, Plan Limits, Recordings

    OpenAI released GPT-Live in July 2026, letting ChatGPT Voice listen and speak at the same time and be interrupted, while search and reasoning are handed to a background model. Based on OpenAI's official documentation, this article covers how the flow of conversation changes, each plan's model and usage limits after the September adjustment, use cases such as cooking and speaking practice, and audio retention and training settings, with a brief note on API pricing.

  • Lifestyle

    OpenAI Discloses Habitat Storage Architecture: The Scale and Limits Behind Over 1 Billion Weekly Users

    On September 11, 2026, OpenAI published an engineering-blog post describing how its storage platform, Habitat, evolved from a Python client library into an independent service, and was rewritten in Rust in the second quarter of this year. This article sets out the request-volume, regional-coverage, and data-volume figures OpenAI states, plus the durability and regional details it omits; this site has not tested any of this and offers no advice on using or buying anything.

  • Lifestyle

    NVIDIA launches DGX Spark 64GB: on sale October 23 from $4,999, two units can be linked into 128GB

    On October 2, 2026, NVIDIA announced a more affordable 64GB memory version of its DGX Spark personal AI computer, available from October 23 through six makers including Acer and ASUS. It is aimed mainly at developers and researchers who want to run AI models on their own machines. Below we summarize the specs NVIDIA published, its claims about linking two units, and what it means for general readers.

Latest travel guides

Sources

Lifestyle