OpenAI’s API deprecation notices dated October 1, 2026 set shutdowns for four speech-model IDs on January 6, 2027 and three text/code IDs on April 1, 2027. Teams using these models in applications, coding agents or narration services should inventory their dependencies now; deprecation is advance notice, not an immediate outage.
Plan separate speech and text/code migrations, with the speech cutover first. For each live dependency, record its owner, replacement, compatibility checks and deployment date.
The seven IDs in the October 1 notices
Affected API model ID | Shutdown date | Recommended replacement |
|---|---|---|
| January 6, 2027 |
|
| January 6, 2027 |
|
| January 6, 2027 |
|
| January 6, 2027 |
|
| April 1, 2027 |
|
| April 1, 2027 |
|
| April 1, 2027 |
|
These are the exact mappings in OpenAI’s schedule, not a list of every upcoming retirement. The notices provide dates but no shutdown time or timezone. Their scope is the OpenAI API; check separate schedules for ChatGPT/Codex subscriptions, Azure deployments and gateway-managed aliases.
Inventory the deployed model, not just the source string
Start with a filename-only repository search. This illustrative command finds candidates for review; it does not change files or prove that every similarly named variant shares a deadline.
rg -l --hidden \
-g '!.git/**' -g '!node_modules/**' -g '!vendor/**' \
'gpt-5\.3-codex|gpt-5\.1|gpt-5\.4-nano|tts-1(-hd)?|gpt-4o-mini-tts' .Inspect deployment configuration, CI jobs, model routers, fallback settings, saved prompts and queued or batch payloads too. Use API request telemetry to identify what actually runs. Separate active dependencies from documentation, historical examples and fixtures before editing.
Include the undated gpt-4o-mini-tts alias: its model reference currently maps it to the December 15 snapshot above. Do not assume an alias will move automatically. Likewise, investigate dated text-model snapshots found by the search without adding unsupported deadline rows.
For every active route, record the owner and workload; provider, endpoint and resolved model; effective reasoning effort; region and service tier; required output format; acceptance criteria; supported fallback; and planned cutover date.
Speech: verify the interface before changing the model
OpenAI recommends gpt-realtime-2.1-mini, but the Speech create reference does not list it among supported models for /v1/audio/speech. That is a documentation boundary, not evidence from an API test. A drop-in audio.speech.create substitution is not established, and the notice does not announce retirement of the entire endpoint.
The documented Realtime workflow uses sessions, typically WebRTC in a browser or WebSocket on a server. For WebSocket output, audio arrives in event chunks through response.output_audio.delta; completion events do not contain the audio bytes. A service that expects an MP3 file needs to account for audio collection and format handling.
Evaluate scripted narration separately from interactive conversation. The following are proposed acceptance checks, not results from testing:
Script fidelity: does the output preserve every required word, number and name, without omissions or added commentary?
Voice and delivery: check pronunciation, language, pacing and the chosen voice against the product’s requirements; voice identity parity has not been established.
Output handling: verify encoding, sample rate, chunk assembly, file delivery, cancellation and partial failures. For interactive products, include interruptions and long sessions.
Keep billing units separate. TTS-1 costs $15 per million input characters and TTS-1 HD costs $30. The Realtime Mini price table instead lists $20 per million audio output tokens, plus separate text and audio input charges. Those numbers alone do not establish which route is cheaper for your scripts.
If the application also converts incoming speech to text, track that dependency separately using the transcription migration guide. Transcription and speech generation are different tasks with different retirement notices.
Text and code: preserve effective settings first
Use the notice’s exact targets as the initial evaluation candidates: Sol for the two coding/general-purpose routes and Luna for the nano route. Choosing GPT-6.1 Sol instead is a separate migration decision; it is not the Sol ID named in this notice.
A hidden change can come from an omitted setting. GPT-5.1 and GPT-5.4 nano default to reasoning effort none, whereas GPT-6 Sol and Luna default to medium. Start evaluations with the current effective effort explicitly set where supported, then tune deliberately.
OpenAI’s migration guidance identifies request changes to check:
For Sol/Luna function calling in Chat Completions,
reasoning_effortmust benone. Use Responses for reasoning with tools.When effort is not
none, removetemperature,top_pandtop_logprobs. Also remove Chat Completionslogprobsor the Responsesmessage.output_text.logprobsinclude entry, as applicable.If used, migrate
prompt_cache_retentiontoprompt_cache_options.ttland review the current cache-write rules.
Compare representative coding, extraction and tool-use cases before rollout. Record correct completion, schema-valid output, tool arguments, retries, end-to-end latency and billed usage. Keep workload roles distinct: passing a coding task does not qualify a high-volume classification route. For longer agent runs, the GPT-6 production guide covers reasoning, compaction and recovery controls.
A worked cost comparison—not measured savings
Assume a workload totals 1 million uncached input tokens and 100,000 billable output tokens, including any reasoning tokens. Each request stays at or below 272,000 input tokens. Use Standard processing, unchanged token counts, and no caching, regional premiums, tool fees or retries. The cells below add input cost and output cost at published USD rates.
Migration and price sources | Before | After | Change |
|---|---|---|---|
$1.75 + $1.40 = $3.15 | $2.00 + $1.00 = $3.00 | −4.8% | |
$1.25 + $1.00 = $2.25 | $2.00 + $1.00 = $3.00 | +33.3% | |
$0.20 + $0.125 = $0.325 | $0.10 + $0.05 = $0.15 | −53.8% |
For Codex to Sol, let I and O be input and billable output in millions. The respective costs are 1.75I + 14O and 2I + 10O. Under these assumptions, Sol becomes cheaper when output exceeds one token per 16 input tokens. Input-heavy workloads can therefore move in the opposite direction from the example.
The Sol and Luna references apply higher rates to the full request above 272,000 input tokens: double input/cache rates and 1.5 times output rates. Recalculate from your actual context lengths, caching and service tier. This arithmetic does not show equal quality or predict production savings.
Close the migration with evidence from live traffic
Set application-specific quality, error, latency and cost thresholds before the canary. Validate the replacement in the actual project and region, including access and rate limits.
Roll out in stages, with an owner able to stop or redirect traffic when thresholds fail. Qualify a supported fallback independently; after shutdown, rollback cannot depend on the retired model.
Complete cutover before the relevant deadline, leaving time to fix failures. Check deployed requests, queued work, scheduled jobs and fallback paths for residual old IDs before closing the migration register.
Methodology: AI-assisted reporting and implementation analysis based on OpenAI documentation checked October 2, 2026. The cost example is arithmetic from published prices. No model calls, voice auditions or production migration tests were performed for this article.
