# Migrate three OpenAI API model IDs before the April 2027 shutdown

> A practical inventory, evaluation, and staged rollout plan for replacing three OpenAI API model IDs before their April 2027 shutdown.
> By Dave · 2026-10-08
> Source: https://otf-kit.dev/blog/openai-api-model-retirement-migration-checklist

OpenAI’s API deprecation notice gives teams using `gpt-5.3-codex`, `gpt-5.1`, or `gpt-5.4-nano` until April 1, 2027, before those model IDs are removed. The listed replacements are `gpt-6-sol` for the first two and `gpt-6-luna` for `gpt-5.4-nano`. Treat that date as a migration deadline, not as a reason to replace every model reference in one pull request.

A safer sequence is: find every configured ID, put the choice behind one application setting, compare the old and recommended models on representative work, then move traffic in stages. This guide focuses on an API application you own. Codex product configuration and other OpenAI products may have different migration paths; check their own notices rather than assuming the API table applies to them.

## Start with an inventory you can repeat

Search the repository, deployment configuration, and operational scripts for the exact identifiers. A source-only search is a useful first pass:

```bash
rg -n --hidden \
  --glob '!node_modules' --glob '!.git' \
  'gpt-5\.3-codex|gpt-5\.1|gpt-5\.4-nano' .
```

Then check the places a repository search cannot see: environment variables in each deployed environment, secret-manager values, scheduled jobs, queued task payloads, admin dashboards, and model names stored in a database. Record the service, environment, owner, current model, and fallback behavior for every hit. A model may be selected by a tenant setting or a persisted workflow version even when the application source has no literal string.

Keep this inventory as a migration artifact. It helps answer the important question after rollout: which production caller still sends a retiring ID? If the API request layer logs a model name already, add a count by model and service before changing traffic. Avoid logging prompts or user content just to observe model selection.

## Make model selection a configuration seam

Centralize the default and replacement IDs so individual routes do not carry their own choices. For example:

```ts
const modelByTask = {
  code: process.env.OPENAI_CODE_MODEL ?? "gpt-5.3-codex",
  general: process.env.OPENAI_GENERAL_MODEL ?? "gpt-5.1",
  small: process.env.OPENAI_SMALL_MODEL ?? "gpt-5.4-nano",
} as const;

export function modelFor(task: keyof typeof modelByTask) {
  return modelByTask[task];
}
```

This keeps the current defaults in one place while allowing a deployment to test a replacement without rebuilding application logic. Before using this pattern, validate environment values at startup against the set of model IDs your application intentionally supports. Fail with a clear configuration error for a typo; silently falling back to another model can hide a bad rollout.

Do not assume the three old IDs map to one replacement. OpenAI’s notice maps `gpt-5.3-codex` and `gpt-5.1` to `gpt-6-sol`, and `gpt-5.4-nano` to `gpt-6-luna`. Preserve the task distinction until your own evaluation shows a reason to change it. Also verify account access and the current model documentation before selecting a replacement in production.

![Dex and Nova show three application workbenches selecting a model through one central configuration console before separate evaluation lanes.](https://cdn.otf-kit.dev/blog/openai-api-model-retirement-migration-checklist/inbody-config-indirection-20261008a.png)

## Build a test set from real work

A replacement is acceptable when it completes the application’s actual job, not merely because a request returns HTTP 200. Choose a small, permissioned set of representative inputs from each task family. Remove personal or customer-identifying details, keep the expected outcome or reviewer rubric, and include cases that exercise the edges your users report.

For code tasks, include changes that need repository context, tests, and a clear patch review. For general responses, include the structured fields or citations your application consumes. For the small-model path, include routine requests where response time and per-request cost matter. Do not compare only easy prompts: include malformed input, ambiguous requests, and cases where the correct answer is to ask for missing information or decline an unsupported action.

Run the same inputs against the current and recommended model with the same system instructions, tools, output constraints, and application settings. If a model’s supported parameters differ, record the adjustment rather than quietly changing the test. Save request IDs, model IDs, timestamps, response status, latency, token usage when returned, and task-specific evaluation results. Keep the prompt data access-controlled; the report should store only what reviewers need.

## Compare behavior, latency, and cost

Set acceptance criteria before looking at the results. Useful checks include task completion, schema validity, code-test pass rate, reviewer corrections, refusal behavior, p50 and p95 latency, and observed cost per completed task. Compare like with like: measure the full application path, including retries and tool calls, and separate first-token latency from total completion time when streaming matters.

Use the actual usage fields and current pricing for your account to calculate cost. Do not infer a price from a model name or multiply a single sample into a monthly estimate without checking request volume and workload mix. A useful unit is cost per successful task, because a cheaper response that needs two repair calls may cost more overall. Keep the raw sample count beside averages; a small test set gives directional evidence, not a guarantee of future performance.

If one replacement changes behavior in a way your rubric rejects, pause that task family. Keep the old model configured while you investigate the prompt, tools, output format, or task routing. Change one factor at a time so you can tell whether the model or application change caused the regression.

![Byte and Luna compare two evaluation benches using output checks, cost, and latency before advancing a small canary cohort.](https://cdn.otf-kit.dev/blog/openai-api-model-retirement-migration-checklist/inbody-testing-staged-rollout-20261008a.png)

## Roll out by task family

Move one task family at a time. Start in a non-production environment with a small, synthetic or approved test set. After it passes, send a limited share of eligible production traffic to the replacement if your architecture supports that safely. Compare the same quality and latency signals, and keep a fast route back to the previous model while the old ID remains available.

Make rollback explicit. Keep the prior model ID in configuration and define the signal that triggers rollback, such as a task-quality threshold, an error-rate increase, or a latency budget breach. Do not use a vague “looks worse” rule after rollout begins; name the reviewer, sample window, and threshold beforehand. Expand traffic only after the first cohort clears those criteria.

When the replacement is stable, remove the old ID from defaults and deployment variables, then search again. Verify queued jobs and stored task settings as well as the checked-in files. Alert on any request that still names a retiring model so a hidden caller cannot survive until the shutdown date unnoticed.

## Keep a dated migration record

Write down the inventory, tested model pairs, evaluation set version, results, reviewer, rollout dates, and rollback decision. Record exceptions with an owner and a due date. That gives the team a concrete way to finish the migration before April 1, 2027, and lets a later incident review distinguish a model change from a prompt or infrastructure change.

If your concern is credential rotation rather than model selection, see [how to set a maximum lifetime for OpenAI project API keys](/blog/openai-api-key-expiration). Model retirement and key expiration are separate controls; completing one does not address the other.

## Sources

- [OpenAI API deprecations](https://developers.openai.com/api/docs/deprecations) — retirement date and recommended replacements for the three model IDs.
