OpenAI published A model guide for the GPT-6 family on October 2, 2026. For teams building API applications, it brings model selection, reasoning settings and long-running agent controls into one guide. This is operational guidance, not a new model launch or price announcement.
The useful production decision is which controls to combine. A workflow that changes reasoning effort has different compaction requirements from one that delegates to hosted subagents. Meanwhile, the application still has to manage pending tools and determine whether the requested work actually finished.
Choose a compatible runtime pattern
The following decision framework combines OpenAI’s reasoning-update rules, compaction guide and Responses Multi-agent limitations. These are alternative starting patterns, not three switches to enable together.
Pattern | When to consider it | Context and effort controls |
|---|---|---|
Fixed-effort single agent | Work is mostly sequential; automatic context management is useful. | Keep effort fixed; configure automatic compaction through |
Adaptive-effort single agent | Routine follow-ups and difficult analysis need different effort levels. | Use standard reasoning mode and |
Hosted Responses Multi-agent beta | Independent, bounded work can run concurrently. | GPT-6.1 Sol is documented as supported. Automatic compaction applies separately to the root and each subagent. |
For adaptive effort, keep the request-level setting unchanged and append an update between responses. Do not combine those updates with automatic compaction or automatic truncation; /responses/compact rejects histories containing them. After explicit compaction with compaction_trigger, reapply the desired update before the next user message. Track effective effort yourself: the returned reasoning.effort still reports the original request-level value. These restrictions are documented here.
Cache preservation means eligibility for reuse, not a guaranteed hit. Prompt caching requires a matching rendered prefix; changing earlier instructions, tools or context can break that match. Put stable material first, then inspect actual usage around effort changes and compaction.
Compaction also serves a different purpose from an audit record. OpenAI’s compaction items carry continuation state in an opaque form. Our design recommendation is to retain an application-owned record of job IDs, approvals and verified results outside that state, so a later run can reconcile what happened.
Then select a model for that pattern
OpenAI’s workload recommendations provide candidates for evaluation, not universal rankings. Filter first for the required endpoint, tools and region; compare the remaining candidates on representative tasks at explicit effort settings.
Model and API identifier | Vendor-recommended starting workload | Uncached input / output |
|---|---|---|
GPT-6 Astra · | The hardest reasoning work | $10 / $50 |
GPT-6.1 Sol · | Complex coding, research and computer use | $2 / $10 |
GPT-6 Luna · | Focused, repeated extraction, classification and summaries | $0.10 / $0.50 |
Prices are US dollars per million tokens at Standard rates, checked October 2, for requests with no more than 272,000 input tokens. The linked model references specify that exceeding that threshold doubles input and cache rates and increases output rates by 50% for the entire request. Cache writes, reads, tools and regional or service-tier premiums need separate accounting; these two columns are not a complete workflow estimate.
Astra and GPT-6.1 Sol do not support none effort. Luna does, and its Chat Completions function calling is limited to that setting. Use Responses for Astra or GPT-6.1 Sol tool calling, and for Luna reasoning with tools. OpenAI’s reasoning documentation also notes that reasoning tokens are billed as output: a short final answer need not be a cheap response.
For launch benchmarks and detailed pricing analysis, see RohitAI’s GPT-6.1 Sol guide. The separate deployment question here is how the chosen model participates in a recoverable workflow.
Separate a response finishing from a job finishing
With async tool calling, mark an application-run function or custom tool with async: true. The model can continue independent work while your application executes it. Return the result using the original call_id; if other turns intervened, continue from the latest response rather than an outdated response ID.
Your application owns the job registry, execution, errors and timeouts. Async calling does not make OpenAI execute your function, and it is not the same as background response generation. Hosted built-in tools and programmatic tool calling are outside this async mechanism.
Track jobs, instructions and verified outcomes separately. A model response can end while a job remains pending; a final sentence is therefore insufficient evidence that all required work is complete.
Illustrative design, not an executed test: consider a coding assistant that starts a test suite while preparing a change summary.
Record the test job and its original call ID before processing any operation that waits on it.
Let the assistant draft the summary while tests run, but keep test-dependent decisions blocked.
Return the actual result on the original call ID. Mark the task accepted only after the application checks the required result; a missing result remains unresolved.
Treat steering as a queued instruction
GPT-6’s mid-turn steering works over Responses WebSockets. The event response.steer.accepted acknowledges that an update is queued—not that behavior has changed. Continue reading events for the successor response. If the API needs a tool result or approval, supply that input through the documented continuation flow without repeating the accepted steering.
Steering does not cancel an already-running tool, undo completed actions or retract output already sent. An instruction to stop changing files therefore needs application-level handling for any write job already dispatched.
Queued steering is connection-local, not stored on the original response. Record the instruction, its steer ID and continuation response ID. After a disconnect, compare the event history with actual job state before replaying anything. The engineering consequence is simple: show “update queued” separately from “update incorporated,” and never use either status as proof of an external action’s outcome.
Use hosted subagents only for separable work
The Responses Multi-agent beta explicitly lists GPT-6.1 Sol and GPT-5.6 models. Do not infer hosted Astra or Luna support from a general reference to GPT-6 orchestration. This feature is also distinct from custom orchestration and the managed Agents API.
Subagents inherit the request’s model and tools. The default concurrency limit is three active subagents across the tree, excluding the root; it does not cap the total number created over the run. Extra agents can add token use, so independent investigations are a stronger fit than workers repeatedly editing the same resource.
This mode does not support /responses/compact, reasoning.summary or max_tool_calls. Enforce run-wide time, spend and action limits in your application. If combining it with async tools, the async compatibility rules prohibit combining those tools with parallel tool calls.
For the broader distinction between model decisions, managed execution and authority, see RohitAI’s Decisions API and agents analysis; its managed-platform settings should not be copied into this Responses configuration.
Buy speed for the part that is actually slow
Separate generation time from tool execution, external waits and human review before paying for a faster service tier. This is a workload-level inference: speeding generation cannot remove time spent waiting for a required external result.
Fast mode shares model rate limits with Standard and can downgrade rapidly ramping traffic. Log the returned service_tier, not just the requested tier. Fast is unavailable with EU data residency for these three models. The current Ultrafast guide establishes broad Astra API availability at low default limits, with US residency or global processing only; it does not establish current GPT-6.1 Sol Ultrafast availability.
Apply the same practical test to interfaces. OpenAI’s October 2 guide recommends direct APIs or connected tools where they can do the step, and computer use where screen interaction is needed. For Responses computer use, the application must keep the browser or desktop environment available: continuing the model conversation does not restore that environment.
A focused pre-deployment check
Build on OpenAI’s deployment checklist with acceptance conditions specific to the runtime pattern you selected:
Compatibility: Record the exact model ID, effective effort, reasoning mode, service tier and tool versions. Confirm the chosen compaction and orchestration controls work together.
Recovery: Exercise a delayed tool result, a failed job and a steering disconnect. Check that required work stays pending until its actual outcome is known.
Economics: Compare cold and warm requests, including cache writes and reads, reasoning output, retries and tool charges. Record end-to-end latency percentiles and cost per accepted task, not just token price.
Completion: Require evidence appropriate to the action—a test result, saved artifact or verified application state. Specify which decisions need human input before the run begins.
For a small workflow, a single agent with automatic compaction is a reasonable starting point. Choose adaptive effort when its benefit justifies explicit context management; choose hosted Multi-agent when genuinely independent work justifies coordination. Evaluate those changes one at a time so any improvement can be attributed to the configuration that produced it.
Methodology: Prepared with AI assistance from OpenAI’s published guide and official API documentation checked on October 2, 2026. The runtime matrix and application-state recommendations are analysis of those interfaces. No models, latency claims or workflows were independently benchmarked; the coding scenario is illustrative.
