OpenAI made GPT-6.1 Sol Ultrafast available on October 8, 2026, in the Responses API and, for eligible plans, Codex and ChatGPT Work. The release gives teams running latency-sensitive coding and agent workflows another speed option, with higher token prices and faster consumption of subscription allowances.
The immediate decision is where shorter waits justify the premium. Keep Standard for work without a tight deadline; compare Fast and Ultrafast on tasks where a person or another process is actually blocked. This is a service-tier rollout for the model released September 29, not a new Sol model launch.
RohitAI’s September 29 Sol guide covers the base model and its launch benchmarks. Its Sol Ultrafast availability statements describe that earlier cutoff; the October 8 documentation supplies the access and pricing details below.
API pricing: six times Standard, three times Fast
The current API rate card lists these US-dollar prices per million tokens for GPT-6.1 Sol requests with no more than 272,000 input tokens:
Service tier | Input | Cached input | Cache writes | Output |
|---|---|---|---|---|
Standard | $2.00 | $0.10 | $2.50 | $10.00 |
Fast | $4.00 | $0.20 | $5.00 | $20.00 |
Ultrafast | $12.00 | $0.60 | $15.00 | $60.00 |
Above 272,000 input tokens, Ultrafast costs $24 input, $1.20 cached input, $30 cache writes and $90 output per million tokens. The model reference applies the higher rates to the entire request, not just tokens above the threshold. Regional processing adds a 10% premium where available.
A worked example: what must the saved time be worth?
Illustrative calculation, not a measured workload: assume 100,000 uncached input tokens, no cache-write tokens and 10,000 total billable output tokens, including any reasoning tokens. Hold token counts and accepted outcomes equal across tiers and use global processing.
Using the rate card above, Standard costs (0.1 × $2) + (0.01 × $10) = $0.30. Fast costs $0.60; Ultrafast costs $1.80. Ultrafast therefore adds $1.50 over Standard, or $1.20 over Fast, per such request.
At an assumed value of $60 per hour of genuinely blocked time, those premiums break even at 90 seconds saved versus Standard or 72 seconds versus Fast. This excludes tool fees, retries, regional premiums, tax and discounts. If someone can do other useful work while waiting, valuing every saved second at their full hourly rate overstates the benefit.
Token-generation speed is only one part of completion time. A planning model is: new task time = unchanged tool and review time + generation time ÷ measured generation speedup. Slow tests, remote tools and human approval do not disappear when tokens arrive faster. No matched Sol Ultrafast latency test was performed for this article.
Codex and Work: access is separate from credit balance
The October 8 product changelog lists Pro $500 and eligible Enterprise/Edu access. Enterprise Ultrafast is disabled by default. The speed documentation says other self-serve plans cannot unlock it merely by purchasing credits.
Eligible Enterprise agreements use credits or USD usage billing; eligible Edu plans use credits. Legacy rate-limit-based Enterprise plans are excluded. Pro $500 uses its included allowance first, then available credits. With an API key, Codex instead follows API token pricing—not the subscription multipliers.
Work and Codex share usage. Their pricing page distinguishes included allowance consumption from paid usage, relative to Standard for the same model:
Mode | Included allowance consumption | Purchased credits / Enterprise pay-as-you-go |
|---|---|---|
Standard | 1× | 1× |
Fast | 2.5× | 2× |
Sol Ultrafast | 8× | 6× |
These are billing and allowance multipliers, not speed measurements. Do not use API dollars to estimate included task counts. Codex credit billing also has no separate cache-write charge; copying the API table into a subscription budget would misstate that bill.
Administrators have two checks: enable the Sol model where required, and enable Ultrafast for permitted Enterprise users or the workspace. Product-surface access and existing per-user spending controls still apply. A saved model or speed default does not grant permission.
Configure the API request, then check capacity and region
Use the existing gpt-6.1-sol model ID with service_tier: "ultrafast". The Ultrafast guide supports Responses over HTTP or WebSockets. This illustrative HTTP request body is not an executed test:
{
"model": "gpt-6.1-sol",
"service_tier": "ultrafast",
"input": "Identify missing test cases in this proposed patch: ..."
}Ultrafast has a separate rate-limit pool from Standard and Fast. OpenAI lists Sol defaults of 1 million tokens per minute for Build, 4 million for Launch and 40 million for Grow. Check the effective limits on your organization before increasing traffic; TPM is capacity, not generation speed. For tier qualification, see the Build, Launch and Grow guide.
Sol Ultrafast supports global processing and US/EU data residency. The current Sol model page also supports Fast with EU residency. In Work and Codex, the speed page specifies US or European inference residency, with Europe defined as the EEA plus Switzerland. Astra’s different regional restrictions should not be applied to Sol.
Residency still needs eligible account and project configuration. OpenAI’s data-controls guide requires approval for abuse-monitoring controls and a Modified Retention amendment for non-US API residency. Inference location does not establish where every external tool or connector processes data.
Choose the tier for the wait it can remove
Keep Standard when completion is not time-critical. Paying more cannot shorten an external test suite or an approval queue.
Evaluate Fast before making Ultrafast the default. At unchanged API token counts, Ultrafast costs three times Fast; compare the additional time saved, not just the difference from Standard.
Reserve Ultrafast for generation-heavy, blocking paths that pass your cost and quality checks. Hold prompts, reasoning effort, tools and cache state comparable; record accepted outcomes, billable tokens, retries and total completion time.
For frequent tool round trips, consider persistent WebSockets. OpenAI recommends them because connection overhead can reduce the benefit of faster generation.
Methodology: AI-assisted reporting and analysis based on published OpenAI documentation checked on October 8, 2026. Calculations are illustrative arithmetic, not benchmarks or customer results. Account-specific access, concurrency and minimum client versions were not independently tested.
