OpenAI Presence Turns Agent Deployment Into Release Engineering

Rohit Ramachandran avatarRohit Ramachandran
OpenAI Presence agent release loop connecting a bounded job, policy, evaluation, deployment, production signals, Codex, and human approval

OpenAI Presence Turns Agent Deployment Into Release Engineering

OpenAI launched Presence without publicly naming a model ID, documenting an API or SDK, or publishing a price card. Those absences explain the product better than another capability list would.

Presence sells the work that begins after a model can hold a convincing conversation. An enterprise picks a bounded job. OpenAI and the customer scope its knowledge and system access, translate policy into permitted actions and approval points, simulate difficult cases, grade the behavior, watch production sessions, investigate failures with Codex, and promote approved changes through a controlled rollout. OpenAI Forward Deployed Engineers and select global systems integrators lead the deployment.

That makes Presence agent operations packaged as a managed product. The agent is only one moving part. In RohitAI's reading, the durable product is the release loop around it: authority, evidence, remediation, and accountable change.

This is a consequential move for OpenAI, but the launch evidence is thinner than the confident branding. OpenAI reports strong results from its own phone-support line, while publishing no traffic volume, issue mix, repeat-contact rate, cost, or independent audit. BBVA and IAG are described as exploring Presence; SoftBank is testing it. Pricing, model routing, service levels, regions, security scope, and export formats remain publicly undisclosed.

The right enterprise question is therefore not, “Can the agent talk naturally?” It is: who owns the evidence and change system that keeps the agent safe, useful, and portable after month six?

The launch is defined by what OpenAI did not ship

The Presence announcement calls the product available today, but the access paragraph is narrower: eligible enterprise customers can enter a limited general availability program, deployments are led by OpenAI FDEs and select global systems integrators, and there is no self-service product yet.

OpenAI also says it will continue supporting voice builders through its frontier models in the API. That distinction matters. Presence is not the next version of the Realtime API, and nothing in the announcement says which model powers it. The public launch page discloses no endpoint, SDK, schema, connector catalogue, context limit, rate limit, hosting design, or supported-region matrix.

There is no Presence entry on OpenAI's business pricing page, either. ChatGPT Enterprise pricing and Realtime API token rates cannot be used as substitutes. A buyer does not yet know whether Presence is billed by conversation, action, successful outcome, audio minute, model usage, deployment labor, or some negotiated bundle.

That is not a small documentation gap. The best current reading is a managed deployment engagement with software inside it, rather than software a developer can independently adopt.

Presence is a governed change loop

The workflow starts with a specific job: resolve a billing issue, support an insurance claim, or handle an employee IT request. OpenAI says the agent receives only the knowledge and system access required for that job. The customer defines what it may do, when approval is required, and when a person should take over.

Before launch, teams simulate common requests, edge cases, and higher-risk scenarios. Graders check the outcome, policy adherence, tool use, and escalation behavior. After launch, production sessions, escalations, and quality signals expose gaps. Codex, using what OpenAI calls a Presence plugin, investigates and suggests changes. Teams compare those proposals with the production version, test them, approve them, and then roll them out under control.

Presence control loop from bounded job through policy, simulation, deployment, production signals, Codex proposals, and human approval

Presence turns agent improvement into a release cycle. Codex proposes; the operating team owns the gate.

This is release engineering in agent clothing.

The useful mental model is not an agent that rewrites itself. It is a versioned system in which policy, prompts, tools, graders, simulations, model choices, and deployment configuration should move together. A failure becomes a regression case. A proposed repair becomes a candidate release. A rollout needs a canary, an alarm, and a rollback path.

That difference is crucial. “Self-improving agent” suggests autonomous mutation in production. OpenAI's own description keeps testing and human approval in the loop. The valuable asset is likely the accumulated regression corpus and change history—not the novelty of a model editing its own instructions.

Benchmark snapshot
Where Fable/Mythos looks strongest
Issues resolved without human assistance
75%
Reduction in human handoffs
15 pp
Reported improvement window
10 days
Self-service availability
No
AreaReported resultWhy it matters
Issues resolved without human assistance
OpenAI phone support
75%Vendor-reported. OpenAI does not publish traffic, issue mix, repeat contacts, false resolution, CSAT, cost, or an independent audit.
Reduction in human handoffs
Codex-assisted loop
15 ppVendor-reported change over ten days. Starting and ending rates, intervention mix, and measurement method are undisclosed.
Reported improvement window
Iteration speed
10 daysEvidence that the operating loop can move quickly, not proof that every change generalized safely or persisted.
Self-service availability
Limited GA
NoEligible enterprises need an account team and an FDE- or integrator-led deployment.

The support line proves operation, not procurement readiness

Presence powers OpenAI's English-language AI phone support. OpenAI says the line met or exceeded undisclosed internal frontline-human quality benchmarks within weeks, now resolves 75% of inbound issues without human assistance, and cut human handoffs by 15 percentage points in ten days.

Those are promising operating claims. They are not yet a reproducible case study.

The current AI phone support Help Center page adds an important boundary. It says the Presence-powered service handles routine product, account, and troubleshooting questions, but cannot submit a report or request, connect a caller to a live agent, initiate an escalation or account review, or guarantee follow-up. That makes “resolution” especially sensitive to definition. A conversation ending without a human is not automatically a correctly closed support issue. The launch reports a 15-percentage-point reduction in human handoffs, while the current public line says it cannot connect callers to a live agent; OpenAI does not explain which historical or internal handoff flow the metric measures.

The same page says calls may be retained to improve services, and that audio clips and transcripts may be used to train models with efforts to reduce identifying information in training data. It also says conversations are disassociated from phone numbers and cannot be retrieved or exported. Those are terms for the public phone line, not evidence about a negotiated enterprise Presence contract. OpenAI's enterprise privacy page does not currently name Presence, while its general Services Agreement says customer content is not used to improve services without explicit agreement and leaves the applicable order form in control. Buyers need a Presence-specific data schedule rather than assumptions borrowed from the public phone line, ChatGPT Enterprise, or the API.

The named enterprise examples do not close the evidence gap. BBVA is a design partner exploring voice support in Mexico. SoftBank is testing Japanese conversations and offers a qualitative endorsement. IAG is exploring support during demand spikes such as severe weather. The announcement does not provide independent production volume, correct-resolution rates, cost, safety results, or ROI for any of the three.

This distinction is not academic. A 2026 field experiment in customer service reported improvements in speed and subjective ratings without an improvement in an objective retrial measure. A separate evaluation-driven deployment paper found that carefully calibrated offline evaluations could correlate with online gains across five production deployments. Neither study evaluates Presence. Together they make the practical point: the evaluation design determines whether an improvement loop learns from business outcomes or merely from convenient proxies.

Most of the feature checklist is already table stakes

Presence combines approved actions, simulations, evaluations, monitoring, guardrails, and handoff. That sounds comprehensive because it is. It is not unique.

Sierra publicly describes voice, chat, email, and WhatsApp agents across 59 languages, system actions, built-in testing, automated monitoring, and expert deployment support. Across Salesforce, Google Cloud, and AWS, the same broad primitives—testing and evaluation, tools, versioning or observability, and escalation—are already public.

The commercial units are different enough that sticker prices are not apples-to-apples. That is precisely why Presence's missing billing unit matters.

PlatformPublic operating modelSelected public economic unitQuestion the unit creates
OpenAI PresenceLimited GA; FDE- or integrator-led deploymentNone disclosedWhat is included: models, telephony, integrations, evals, FDE work, and ongoing improvement?
SierraManaged agent platform with an expert deployment teamAgreed outcomes; blended consumption can applyHow is a successful outcome defined, attributed, disputed, and repriced?
Salesforce AgentforceCRM-native agent platform with builder and service surfaces$500 per 100,000 Flex Credits or $2 per customer-facing conversation bought in advanceHow many metered actions does one correct resolution consume, and what other Salesforce costs apply?
Google Conversational AgentsFlows and playbooks with per-request and per-audio-second billing$0.007–$0.012 per chat request; $0.001–$0.002 per billed voice secondHow many requests and billed audio seconds does the complete workflow require?
Amazon Connect CustomerContact-center channel plus bundled agent design, testing, and observability$0.010 per chat message sent or received; $0.038 per voice-service minute plus communicationsWhich telephony, external-model, knowledge, gateway, simulation, and third-party charges sit outside the base unit?

Sources: Sierra outcome pricing, Salesforce Agentforce pricing, Google Conversational Agents pricing, and Amazon Connect Customer pricing. Sierra says unresolved conversations and escalations are uncharged in most cases, while allowing blended arrangements. Salesforce's Flex Credit and conversation routes cannot coexist in the same org. A Google request is an API call, and one turn can require several; billed voice time is not the same as wall-clock call time. Amazon Connect Customer's charge applies across channel use in the account, while communications and some external services can add cost. None of these selected units is an end-to-end resolution price.

RohitAI's read is that the billing unit exposes where each vendor places commercial accountability. Sierra ties payment to an attributed business result. Salesforce can meter actions through Flex Credits or sell conversations and per-user access. Google meters platform requests and billed audio seconds. AWS bundles much of the AI layer into a contact-center channel charge, with communications and some external services added. Until OpenAI discloses its unit, no buyer can tell whether Presence is priced like software, labor, inference, an outsourced outcome, or all four.

Presence's possible moat is the organization around the loop

Voice is an effective wedge because its failures are immediate and expensive. But voice quality itself is unlikely to be the durable moat. OpenAI, Google, specialist audio labs, and orchestration vendors are all pushing latency, turn-taking, entity capture, tool use, and telephony forward.

RohitAI's earlier analysis of GPT-Realtime-2.1 argued that the value in production voice moves toward monitoring, compliance, handoff, analytics, and tuning around the model. Presence is OpenAI moving directly into that surrounding layer.

The differentiated bet is organizational:

frontier models
      +
production traces and quality labels
      +
Codex-assisted investigation
      +
forward-deployed engineers
      +
customer policy owners and approval gates
      =
a managed improvement system

Most visible features are copyable. The harder bet is a loop that connects the model provider's research teams, deployment specialists, code agent, production telemetry, and customer approvals. If OpenAI can shorten the path from a real failure to a safe release without hiding the evidence from the customer, Presence can earn its place.

The FDE- and integrator-led delivery model is also the constraint. Embedded experts raise the chance that a difficult integration works. They can create a human-capacity ceiling, add bespoke implementation cost, and leave a handoff problem: who owns the eval suite, policy changes, broken integrations, incident response, and rollback if or when the launch team leaves?

Limited GA can therefore serve as managed product discovery. Repeated deployments could show OpenAI which policies, failure modes, tools, and workflows are reusable enough to become modules. OpenAI explicitly says generalized deployment insights feed research and product development. The contract still needs to define what “generalized” means, how customer material is separated, and whether customers can opt out.

Lock-in moves above the model

Enterprise AI procurement often asks whether a vendor can swap one model for another. Presence exposes a more important portability problem.

After a year in production, the valuable system may include hundreds of policy rules, tool schemas, simulation cases, grader definitions, red-team prompts, annotated traces, escalation labels, version histories, Codex change proposals, and undocumented knowledge held by the deployment team. A customer can switch the underlying model and still be unable to move the operation.

Operating assetWhy it compoundsProcurement test
Policies and action schemasEncode what the agent may know, decide, prepare, and executeExport them in documented, versioned formats and replay them outside Presence
Eval and simulation corpusCaptures edge cases, prohibited outcomes, regressions, and business definitions of qualityRun the same held-out suite against a replacement stack before contract renewal
Traces and quality labelsShow how real sessions fail and which repairs actually workExport complete events, tool calls, denials, annotations, and timestamps at useful scale
Version and change historyLinks production outcomes to prompts, policies, tools, graders, and model routesReconstruct any deployed release and preserve rollback evidence after exit
Deployment know-howLives partly in FDE and integrator decisions that may never reach documentationSet a post-deployment RACI, documentation standard, and knowledge-transfer acceptance test

This is the same broad shift RohitAI identified when AWS made the agent session a managed cloud resource: the control plane becomes harder to move than the model call. Presence may create an even stickier asset because it combines software state with services knowledge.

A plausible OpenAI stack, with an important asterisk

OpenAI's recent enterprise moves fit together neatly on paper.

At the bottom are frontier models and APIs. OpenAI Frontier is described as a horizontal platform for enterprise context, tools, identity, permissions, evaluation, and governance across OpenAI, customer-built, and third-party agents. Presence looks like a packaged, conversation-focused workload above that control layer. The OpenAI Deployment Company represents adjacent FDE capacity and a partner model that could support complex deployments.

That four-layer map—models, Frontier, Presence, Deployment Company—is analytically useful. It is not a confirmed architecture. The Presence announcement does not say it runs on Frontier or formally belongs to the Deployment Company. Buyers should not design around an integration OpenAI has not documented.

Still, the direction is clear. Earlier RohitAI coverage described GPT-5.6, Codex, ChatGPT Work, and the API as a product stack. Presence advances that strategy from model and orchestration surfaces into hands-on operational ownership. RohitAI reads OpenAI as competing for the layer where enterprise behavior is specified, measured, and changed.

Three sensible adoption paths

Managed
Use Presence for one bounded job

Best for a high-volume voice or chat workflow where integration, policy, evaluation, and ongoing tuning are harder than the model call—and where the organization accepts a services-led limited-GA program.

Build
Keep the control plane in-house

Best when self-service APIs, model choice, custom infrastructure, precise data controls, or exportability matter more than speed from an embedded OpenAI team. Continue evaluating the Realtime API and other composable stacks separately.

Hybrid
Rent execution, own the evidence

Use Presence for a narrow production lane while keeping policy sources, tool contracts, eval cases, outcome labels, and release records in customer-controlled systems that can test another provider.

The hybrid path is the most defensible default. It lets a company benefit from OpenAI's deployment loop without making the loop's evidence impossible to recover.

The acceptance test should look like release engineering

A polished demo should not pass procurement. A reproducible operating contract should.

Presence pilot and procurement checklist
01Define one bounded job, its prohibited outcomes, and the exact events that require fresh approval or human takeover
02Request denominators behind the 75% and 15-percentage-point claims: traffic, issue mix, baseline, repeat contacts, CSAT, latency, cost, and review method
03Document the underlying model and routing policy, supported regions and languages, channels, concurrency, latency targets, uptime, and recovery design
04Version policies, prompts, action schemas, graders, simulations, model routes, and deployment configuration as one release
05Treat every Codex proposal as an untrusted candidate change: held-out evals, separation of duties, canary rollout, alarms, and one-step rollback
06Use short-lived job-scoped credentials and two-phase commit for irreversible actions; log attempted, denied, prepared, approved, and executed operations
07Add session-level controls for cumulative spend, action counts, risk budgets, context changes, and re-authentication—not only per-tool allowlists
08Design handoff as a product path: triggers, queue capacity, identity state, context transfer, transcript summary, and behavior while waiting
09Negotiate retention, improvement use, de-identification, residency, integrator access, deletion, legal hold, and incident response specifically for Presence
10Export policies, tools, evals, simulations, traces, labels, versions, and change proposals during the pilot and prove they can be replayed elsewhere
11Assign post-FDE ownership for policy changes, integration breakage, eval maintenance, incidents, cost anomalies, approvals, and rollback
12Compare total cost per correctly resolved workflow against an API-built OpenAI stack and at least one independent agent platform

The safety model must also extend across a session. A sequence of individually permitted tool calls can accumulate into an unsafe outcome. RohitAI's analysis of OpenAI's long-horizon agent failures explains why the session—not one tool invocation—needs cumulative limits, trajectory review, pause, and commit gates.

For banking and insurance, support flows must not cross into fully automated high-stakes decision-making. OpenAI's Usage Policies disallow automating high-stakes decisions without human review in areas including financial activities, credit, and insurance. A Presence agent can help explain or gather information without receiving unchecked authority to adjudicate a claim or approve credit.

Measure the outcome the customer would recognize

The cleanest pilot metric is not containment. It is a stricter definition of correct resolution:

correctly resolved workflow
= intended outcome completed
+ policy followed
+ identity and approvals verified
+ no unauthorized side effect
+ no repeat contact inside the agreed window
+ evidence sufficient for audit

Track that alongside false resolution, appropriate escalation, unauthorized-action attempts, tool-error rate, p50 and p95 latency, customer satisfaction, human handling time, and total cost. Segment every number by issue type and risk tier. A 90% result on password-reset questions must not hide a 40% result on billing disputes.

For every proposed improvement, keep a control group or at least a stable held-out set. A change that reduces handoffs can do so by solving more problems, by becoming overconfident, or by silently narrowing what counts as escalation. Only downstream evidence separates those outcomes.

Cost needs the same discipline:

total cost per accepted resolution
= platform and model use
+ telephony and messaging
+ tool and data services
+ FDE and integrator work
+ human escalation
+ quality review
+ change management
+ cost of repeat contacts and remediation

That denominator lets Presence compete fairly with a cheaper API stack that consumes more internal engineering, or an outcome-priced vendor whose definition of success differs from yours.

The RohitAI read: Presence is services becoming software

Presence matters because OpenAI has acknowledged, in product form, what experienced agent teams already know: the model is not the production system. Policies, permissions, evals, telemetry, escalation, and controlled change are the system.

The company's strongest advantage is not that it has invented those components. It has not. The advantage is the chance to join them into a fast loop with model researchers, Codex, FDEs, integrators, and customer operators. If that loop produces safer releases faster than an enterprise can build them, Presence can be valuable even while every visible feature has a competitor.

The risk is equally clear. The same loop can concentrate the customer's policy, evidence, operational memory, and improvement process inside OpenAI. Model lock-in is easy to discuss because model IDs are visible. Eval lock-in, trace lock-in, and FDE-knowledge lock-in arrive quietly.

My expectation is that OpenAI will spend the next year turning bespoke deployments into reusable modules, then expose more programmatic and self-service surfaces once the patterns stabilize. Pricing will probably remain custom or services-bundled until OpenAI knows whether the repeatable unit is an action, conversation, accepted outcome, or managed workflow. Portability of policies and evidence will become a serious buying criterion before model portability does.

Presence is therefore an important launch with an incomplete buying case. It deserves attention because it moves OpenAI up the stack. It deserves scrutiny because the public proof, economics, data scope, and exit path are not yet detailed enough to accept on branding alone.

FAQ

Is there a public OpenAI Presence API or SDK?

No public Presence API or SDK has been announced or documented. OpenAI describes Presence as a deployed product in limited general availability for eligible enterprise customers. Deployments are led by FDEs and select global systems integrators, and the product is not self-service. OpenAI says voice builders will continue to have frontier-model access through the API.

Which model powers Presence?

OpenAI has not said publicly. The launch names no model or routing policy. Do not assume Presence uses GPT-Realtime-2.1, GPT-5.6, or any particular API price card.

Does Presence autonomously improve itself?

Not according to the public workflow. Codex investigates signals and proposes updates. Teams test those proposals against the production version, approve them, and control rollout. That is a human-governed release loop.

How much does Presence cost?

OpenAI has not published a Presence price or billing unit. Interested organizations must contact their account team. Buyers should ask separately about model use, telephony, integrations, FDE labor, systems-integrator work, evaluation, and ongoing support.

Is Presence the same as OpenAI Frontier?

No confirmed relationship has been published. Frontier is positioned as a horizontal enterprise platform for context, tools, identity, permissions, evaluation, and governance across many kinds of agents. Presence is a managed product for bounded voice and chat jobs. They may fit together, but OpenAI has not documented that architecture.

What is the most important pilot metric?

Cost per correctly resolved workflow, with repeat contacts, false resolution, policy adherence, escalation quality, latency, and human effort included. Containment alone is too easy to improve without improving the customer's outcome.