OpenAI’s Python SDK v3.27.0, published on October 9, 2026, adds support for prewarmed hosted environments. For Agents API builders, the change separates environment preparation from session creation: prepare compute first, then attach a session by environment ID. The create endpoint requires prewarming-beta access; upgrading the library alone does not grant it.
The useful question is whether your application knows enough about an upcoming task to prepare its environment before the user needs it. That can move setup off the user’s waiting path, but the reviewed sources do not establish a measured speedup or a prewarming-specific billing guarantee.
When preparing ahead can help
Consider prewarming for predictable jobs or interactive workflows with an early signal of demand. Keep session-first provisioning as the baseline for irregular traffic until the benefit of earlier setup outweighs unused capacity and lifecycle management. This is an implementation judgment, not an OpenAI performance claim.
A simple analytical model makes the tradeoff explicit. Let P be preparation time, L the lead time before demand and A the remaining session/attachment overhead. Assuming unchanged preparation and overhead, no expiry and no queueing, the infrastructure wait changes from P + A to max(0, P − L) + A. The potential preparation wait removed is min(P, L). Model inference and tool execution still take time; these variables are not measured results.
The lifecycle is: prepare the environment, check readiness, then attach the session. OpenAI’s create and session-create contracts support those separate steps.
1. Confirm access and create an environment
This guide uses openai==3.27.0, the Python API client—not the separate openai-agents orchestration package. The Agents API quickstart requires OpenAI-Beta: agents=v1, which the SDK supplies automatically. It also specifies api.agents.read and api.agents.write permissions, plus api.responses.write for inference. Confirm prewarming access for the project you will actually use.
The following documentation-derived sketch has not been executed. It assumes an already configured client in your trusted application and an opaque workload_key saved before the request. Networking is explicitly disabled, so this minimal configuration assumes no package downloads or outbound task dependencies. The method and parameter names come from the pinned Python client.
environment = client.beta.agents.environments.create(
environment={
"type": "openai_hosted",
"network": {"access": "disabled"},
},
idempotency_key=workload_key,
)
environment_id = environment.id
# Persist this ID with the workload key and exact request parameters.For a real workload, choose its packages, files, setup commands and network policy during preparation. The prewarm configuration type supports those settings and reusable templates. Keep the intended configuration in your own durable record; do not select an environment just because its status is ready.
OpenAI’s creation-idempotency contract retains keys for 24 hours within the authenticated organization, project and creator. The same key and JSON parameters return the original environment’s current state. Changed parameters or an incomplete hosted creation return 409; without a key, each request creates another environment. A key reused after retention expires may create a new one.
Implementation recommendation: persist the creation intent before sending it, keep retries tied to identical parameters and investigate conflicts rather than replacing the key reflexively. Save the returned environment ID as soon as it is available.
For inventory reconciliation, client.beta.agents.environments.list(limit=100, order="desc") lists the authenticated principal’s hosted environments. The list reference documents pages of 1–100 entries, a default of 20 and an after cursor. Follow pagination and inspect statuses locally; there is no documented readiness filter, and listing is not a reservation.
2. Treat webhooks as signals to check current state
The new agent.environment.ready event reports setup completion before session attachment; agent.environment.failed reports setup failure at that stage. In both, data.id is the environment ID, while the top-level id identifies the event. The failed-event type does not include a detailed error object.
Follow the webhook guide: verify signatures using the original payload, return a prompt 2xx response and handle duplicate deliveries. OpenAI documents delivery retries for up to 72 hours and a webhook-id header for deduplication. A practical design is to durably enqueue the verified event before acknowledging it, then let a worker reconcile the environment.
Recommended handling for an environment awaiting its first attachment: retrieve the recorded ID through the environment-retrieve endpoint and apply the following checks. These are application decisions, not additional server-side guarantees.
Current result | Recommended action |
|---|---|
pending | Wait with a deadline and bounded retries; do not attach yet. |
ready | Check workload/configuration ownership and acquire a durable allocation claim before attachment. |
failed or expired | Stop this allocation attempt and inspect its history before deciding whether to create a replacement. |
connected, disconnected, suspended or unfamiliar | Reconcile the existing ownership/session state; do not treat it as unused ready capacity. |
The v3.27.0 Python status type lists pending, ready, connected, disconnected, expired and failed; the current retrieve reference also lists suspended. Handle unfamiliar values explicitly instead of assuming the pinned client describes every server state.
The retry windows differ. Creation deduplication lasts 24 hours, but webhook delivery can continue for 72. The design implication is that a late event should resolve an existing workload/environment record—not blindly repeat creation under an old key. If an event arrives before the creation response has been recorded, retain it for reconciliation rather than allocating again.
3. Attach by ID without configuration overrides
The session-create contract forbids combining environment_id with a template or inline configuration. Do not add packages, files or network overrides to the attachment request. Choose a prepared environment that already matches the workload.
This second unexecuted sketch assumes environment_id is recorded, saved_agent_id identifies an existing configured agent, and your worker already holds a durable allocation claim. Such a claim is an application safeguard against competing workers, not a documented OpenAI reservation API.
current = client.beta.agents.environments.retrieve(environment_id)
if current.status != "ready":
raise RuntimeError("Environment is not ready for attachment")
session = client.beta.agents.sessions.create(
agent_id=saved_agent_id,
environment={
"type": "openai_hosted",
"environment_id": environment_id,
},
)
# Persist the session/environment mapping before submitting work.A non-streaming hosted session can omit initial input, as shown in the session parameter type. This sketch stops at attachment; it does not submit a task. Record attachment outcomes durably and reconcile ambiguous failures before retrying. Creation idempotency and webhook deduplication do not, by themselves, make task execution exactly once.
Partition prepared inventory by configuration and data-isolation boundary. OpenAI’s sandbox-security guidance calls for separate environments where users or workloads must not share data, and keeping application credentials outside agent compute. A generic pool should not erase those boundaries.
What to establish before relying on a warm pool
Lifetime and cleanup. The general hosted guide describes possible deletion after an hour without activity or keep-alives and session deletion for cleanup. That is not a documented lease for an unattached prewarmed environment. The reviewed Python environment resource exposes no standalone delete method. Establish expiry, cleanup, inventory quotas and attachment/reuse rules before speculative allocation.
Actual charges. The standard container price table lists, for example, $0.12 for a 4 GB container per 20-minute session, with per-minute billing and a five-minute minimum for eligible sessions. It does not settle when an unattached prewarm starts billing, idle charges or failed-setup charges. Do not turn that list price into a promised warm-pool budget.
Configuration and file limits. The hosted guide documents session provisioning sizes, but
container_sizeis absent from the reviewed prewarm-create type. Do not copy session sizing fields into prewarm requests without confirmation. The file guide lists 5 MiB per inline file, 10 MiB total inline data per creation request, 50 MiB per Files API copy and 200 MiB per published artifact; check the path your workload uses.
For a pilot, compare session-first and prewarmed flows using identical configuration and inputs. Record preparation, readiness, attachment, first-tool and completion times alongside billed usage. Exercise failed setup, duplicate and delayed events, competing workers and stale IDs. Confirm the final task output—not just readiness. These are proposed checks, not tests performed for this article.
Start with one predictable workload and a bounded inventory. Scale only after its measured waiting-time reduction, unused-capacity cost and recovery behavior justify the extra machinery. Model throughput and account spending controls still need separate planning; the related OpenAI Build, Launch and Grow usage-tier guide covers that question, not prewarming eligibility.
Methodology: AI-assisted reporting and implementation analysis based on OpenAI’s release metadata, pinned SDK source and official documentation checked on October 9, 2026. No live Agents API calls, account-access checks, latency benchmarks or billing experiments were performed.
