OpenAI Decisions API and Computer-Using Agents: A Stack for Bounded Autonomy

Rohit Ramachandran avatarRohit Ramachandran
Bounded autonomy stack separating finite decisions, agent orchestration, computer use, runtime ownership, and private inference

OpenAI Decisions API and Computer-Using Agents: A Stack for Bounded Autonomy

OpenAI used DevDay 2026 to introduce two products that look unrelated at first glance. The Decisions API chooses from a finite set of answers. The expanded Agents API can now operate an OpenAI-hosted browser, invoke tools, preserve a working session, and continue a longer task.

One narrows a model’s freedom. The other gives a model more ways to act.

Put them together carefully and you get a useful production pattern: bounded autonomy. A fast decision layer selects an allowed route. An agent handles the work that genuinely needs reasoning. A separate policy layer controls credentials, network access, consequential actions, and recovery. An application-owned record decides whether anything important actually happened.

Put them together carelessly and you get a demo that can classify a request, browse a website, and leave your team unable to answer a basic incident question: Who authorized that side effect?

This follows RohitAI’s OpenAI Agents API and Codex harness guide. That covered the runtime; this covers its control system, authority, recovery, and privacy boundaries.

The point is not to make every workflow autonomous. It is to give each part of the workflow exactly as much freedom as it needs—and no more.

Availability snapshot — September 29, 2026, 19:22 UTC. OpenAI announced the Decisions API in limited preview, with broader release planned in the following days. Computer use is documented as available in the Agents API. The recap lists Agents API access in Codex and ChatGPT Work for Pro 500 and Enterprise. Verify changing limits, regions, and plan eligibility before committing an architecture.

The fact layer: what OpenAI actually released

OpenAI’s DevDay 2026 recap describes more than twenty announcements. Four are directly relevant to this stack.

First, the Decisions API accepts text or images as context and chooses among answers supplied by the developer. OpenAI positions it for jobs such as classification, routing, and selecting a next action. The announcement says it uses Luna. At the timestamp above, OpenAI had announced a limited preview but had not published a stable public reference page with an endpoint, request schema, pricing table, or service-level commitment that we could verify. Any code sample pretending otherwise would be fiction, so this article does not include one.

Second, the Agents API now includes computer use. OpenAI’s Agents API overview describes a managed Codex harness that handles session state, model orchestration, compaction, and recovery while the application supplies tools and an execution environment.

Third, the new computer-use guide documents an OpenAI-hosted browser. An application starts a session, follows task events, answers origin-approval and sign-in requests, and reviews the resulting activity.

Fourth, Zero Data Retention with Private Safety Processing is available for eligible API traffic, while OpenAI says a preview of Private Inference is coming this fall. Those controls do not automatically apply to every endpoint.

Here is the important separation between verified facts and our interpretation:

LayerVerified product roleWhat it does not prove
Decisions APISelects from developer-defined finite answers using text or image contextThat the selected answer is fair, calibrated, or safe to execute
Agents APIRuns a managed agent harness with sessions, tools, orchestration, and compactionThat your business operation is exactly-once or recoverable
Computer useLets an agent work through a hosted browser, subject to origin and sign-in handlingThat every consequential click receives a human confirmation
Private Safety ProcessingSupports automated safety review for eligible Zero Data Retention traffic without OpenAI retaining prompts and responsesThat Agents API sessions are ZDR-eligible

The architecture: five jobs, five boundaries

The easiest agent stack to demo is a loop: give a model a goal, connect a browser, and keep going until it says it is done. The production version needs more structure.

Bounded autonomy control stackA workflow moves from untrusted context through a finite decision gate, an agent planner, an authority gate, tools or browser, and an independent verifier connected to an application-owned ledger.Bounded autonomy is a chain of custodyEach layer grants a smaller, explicit kind of freedom.ContextFinite decisionPlannerAuthority gateActionuntrusted inputroute or abstainproposepolicy + consenttool or browser→→→→Verifier + application-owned ledgerrecord intent, approval, side effect, receipt, and final business stateagent session state is evidence—not the source of truth

The model may recommend and act, but the application keeps the authoritative record of intent, consent, and outcome.

This diagram contains the thesis of the article. Intelligence and authority should not live in the same undifferentiated loop.

Decisions API: less output freedom, not less accountability

A finite answer set is valuable because it changes the interface contract. Instead of asking a model to produce any string, the application defines the only acceptable choices: refund, request_evidence, escalate, or abstain, for example. Downstream code receives a known category rather than prose it has to parse.

That removes one class of failure. It does not remove the important ones.

A model can confidently choose the wrong allowed answer, miss a rare pattern, or route an unfamiliar case into the nearest familiar bucket. Without an escape hatch, a bad choice can still look valid.

For a serious deployment, evaluate the decision as a decision—not as a language-model response:

  • Define an explicit abstain or human-review option.
  • Measure false-positive and false-negative costs separately. A missed urgent case and an unnecessary escalation are not equivalent.
  • Slice results by language, image quality, product line, geography, and other relevant cohorts.
  • Compare the model with the current rule system and human baseline.
  • Track drift after policy, product, and traffic changes.
  • Keep the chosen answer and the execution permission as separate objects.

The last point matters. A classifier may decide that a request belongs in refund_eligible. That does not mean the model should issue money. Eligibility is a judgment. Issuing the refund is a side effect with limits, authorization, and a receipt.

Computer use: an origin gate is not an action gate

The Agents API computer-use workflow is more explicit than many browser-agent demos. The application creates a browser-backed session, follows the stream of task events, handles required origin approvals or sign-in, waits for completion, and reviews the activity. OpenAI’s API changelog notes that website-origin approvals and sign-in are handled by the integrating application.

Every new website origin requires approval, including a public site. That is a meaningful boundary: a page cannot silently send the browser anywhere it likes. It is also easy to misunderstand.

OpenAI’s guide explicitly warns that origin approval does not guarantee confirmation before each consequential action. Once an origin is approved, the browser may be able to click a purchase button, submit a form, or delete something available on that origin. If your product requires a guaranteed checkpoint before those actions, OpenAI recommends constraining the browser to resources that cannot perform them or using a browser runtime you control.

This produces a clean distinction:

Origin approval: “This browser may communicate with example.com.”
Action approval: “This user approved order 8472 for $126.40.”

You need both when the action matters.

Website content must remain untrusted after origin approval. A page can contain prompt injection, misleading controls, or instructions designed to expand the agent’s authority. Network permission does not turn web content into policy. Credentials should be scoped, short-lived where possible, and kept out of prompts and logs.

Two protocol details matter. An HTTP 202 means an approval decision was accepted, not that navigation completed. After a disconnect, fetch the same session and current required actions; blindly replaying work can duplicate it.

Durable agent session, fragile business transaction

Managed sessions are useful. The harness can preserve context, compact a long conversation, and help a task resume. But this is not the same as transactional durability.

Imagine an agent submitting a supplier payment form. The browser receives confirmation, but the event stream disconnects before your service records it. If you restart the task, the agent may submit twice. If you assume the payment failed, your ledger becomes wrong. If you trust the agent’s final sentence, you have replaced accounting with narration.

The application needs an intent and outcome ledger outside the agent session:

intent_id -> requested action -> policy result -> human approval
          -> idempotency key -> tool/browser evidence -> verified outcome

Use typed APIs for money movement, permission changes, deletions, messages, and other high-impact effects whenever one exists. Typed tools can enforce schemas, idempotency keys, amount limits, and approval references. Let the browser gather information or handle low-risk interfaces. It should not become the default transaction bus just because it is visually flexible.

This is the third key separation: agent recovery restores cognition; business recovery restores truth.

Three deployment patterns, and where each breaks

High volume
Decision-first routing

Use a finite decision set for triage, moderation queues, document routing, or selecting a workflow. Escalate ambiguous or costly cases to a reasoning agent or person. Best when choices are stable and actions remain application-controlled.

Knowledge work
Agent with typed tools

Give an agent a narrow tool catalog, explicit schemas, least-privilege credentials, and confirmation gates. This is the safest default for workflows with real side effects because the application can validate every operation.

Legacy interface
Agent with browser access

Use computer control when no reliable API exists or the task is inherently visual. Restrict origins and accounts, separate observation from mutation, and verify outcomes independently. Flexibility is high; guarantees are weaker.

The choice is not about which product looks most advanced. It is about the narrowest primitive that can complete the task.

A common good stack is decision-first, tool-second, browser-last. The Decisions API handles repetitive judgment. A reasoning agent takes the uncertain tail. Typed tools handle important actions. Computer use fills the legacy gaps. Humans approve exceptional risk rather than every mundane step.

Privacy is an endpoint property, not a brand promise

DevDay’s Private Intelligence story is directionally important, but builders need to read the eligibility details.

OpenAI’s Agents API overview currently states that the product retains session state, lets customers delete sessions and published artifacts, currently supports data residency only in the United States, and does not support Zero Data Retention. Using your own sandbox changes where commands run; it does not make the managed harness or model inference ZDR-eligible.

Private Safety Processing solves a different problem. For eligible ZDR API traffic, automated abuse review can run in a hardware-attested environment without OpenAI retaining prompts and responses. Customer-controlled storage holds encrypted review material; detailed results have a documented 30-day time to live. It is configured per project.

Those are concrete controls. They are not an umbrella over every API.

This leads to an uncomfortable but useful architecture rule: if strict ZDR is non-negotiable, an Agents API workflow may be the wrong runtime today, even if another OpenAI API workflow is eligible for ZDR plus Private Safety Processing. Check OpenAI’s current data controls documentation, obtain OpenAI’s prior approval and accept the additional requirements, and evaluate each endpoint, tool, region, and storage layer separately.

Bedrock Managed Agents changes the operator, not the need for controls

OpenAI also documents Bedrock Managed Agents, an adaptation of core Agents API concepts for AWS. In that configuration, the agent loop and model inference run through Amazon Bedrock, while commands and tools can run in AgentCore Runtime or self-hosted compute. Authentication uses AWS IAM and SigV4 rather than an OpenAI API key.

That may fit organizations standardized on AWS identity and networking. It does not promise identical APIs or feature parity. Verify regions, models, tools, limits, and pricing for the selected platform.

This extends the earlier RohitAI analysis of the AWS Bedrock AgentCore managed-agent stack: portability lives at the level of architecture concepts, not necessarily request contracts. A session, sandbox, tool, and approval can exist in both stacks while behaving differently in the details that determine security and operations.

For governance across projects, the earlier OpenAI Terraform provider and API platform guide is the useful companion. Agent safety depends on boring controls—project boundaries, service accounts, budget limits, logging, and policy-as-code—long before a model sees a prompt.

A production checklist for bounded autonomy

Before letting the workflow act
01Write the allowed decisions, an abstain path, and the cost of each error type before selecting a model
02Separate a classification result from the authorization object that permits an action
03Prefer typed tools for high-impact side effects; reserve browser control for observation and legacy gaps
04Require action-level confirmation for purchases, deletions, messages, permission changes, and regulated submissions
05Scope origins, credentials, accounts, tool parameters, spend, and time; treat all website content as untrusted
06Assign an application intent ID and idempotency key before work begins
07Persist approvals, tool inputs, receipts, screenshots or event evidence, and independently verified outcomes
08Resume the same session after disconnects; do not blindly replay tasks or approvals
09Run offline evals, cohort slices, adversarial pages, failure injection, and human-review tests
10Confirm current retention, residency, ZDR eligibility, deletion behavior, plan access, limits, and pricing for every component

For model-side evaluation, OpenAI’s evaluation best-practices guide is a useful starting point. For enforcement, the guardrails and approvals guide makes another important point: lower-level Responses API or SDK workflows do not automatically inherit Codex product protections. Your application owns the gate.

Four implications builders should take seriously

1. The best agent product may have fewer agent decisions

Teams often measure progress by how much of the loop the model controls. A mature system converts repeated judgment into a bounded decision, repeated action into a typed tool, and repeated recovery into workflow code. The agent handles ambiguity.

As the product improves, the autonomous surface can shrink even while automation rises.

2. Approval needs a hierarchy

One “approve” button cannot express network access, data disclosure, financial authority, and acceptance of a result. A useful approval object records the actor, action, resource, limits, expiry, and intent ID.

Expect agent platforms to evolve from approval prompts into something closer to cloud IAM for temporary actions.

3. The audit trail becomes part of the user experience

For consequential agents, users need a receipt: what was decided, which tools were used, what required approval, what changed, and how to undo it. The audit layer becomes the trust interface.

4. Privacy will push orchestration into multiple runtimes

One workflow may use a ZDR-eligible API for sensitive classification, a managed session for research, and a customer-controlled runtime for privileged action. This matches real retention, region, credential, and risk boundaries.

This is why the question “Which agent platform are we using?” will become less useful than “Which runtime owns each stage?”

Frequently asked questions

Is the Decisions API generally available?

Not at the availability timestamp in this article. OpenAI’s DevDay recap described a limited preview on September 29, 2026 and said broader release was planned in the following days. Verify the current product page before relying on access, and do not assume an endpoint, model ID, price, quota, or SLA until OpenAI documents it.

Is the Decisions API just structured output?

The public announcement describes a product optimized to choose among a predefined set of answers using text or image context. That resembles a constrained decision interface, but OpenAI had not yet published enough reference detail for us to claim how its internals, schemas, or guarantees differ from every structured-output pattern. The production value is the contract: finite choices that can be evaluated as decisions.

Does origin approval make browser use safe?

No. It controls which website origin the hosted browser may contact. OpenAI explicitly says it does not guarantee a human confirmation before every consequential action on that origin. Add your own action-level policy and approvals, or use a controlled browser environment that cannot perform the risky operation without them.

Does the Agents API support Zero Data Retention?

The current Agents API overview says no. It retains session state, supports deletion of sessions and published artifacts, and currently supports data residency only in the United States. Recheck the official page because product capabilities can change.

Should I use Agents API, Agents SDK, or Responses API?

Use the Agents API when you want an OpenAI-managed harness and saved session state. Use the Agents SDK when you want more application-level control over orchestration. Use the Responses API for lower-level model and tool primitives. OpenAI’s agent-building overview compares the three; your retention, runtime, observability, and control requirements should decide.

When should an agent use a browser instead of an API tool?

Use a typed API tool when it exists and the action matters. Use computer control for visual tasks, research, or legacy systems that lack a suitable API. For high-impact browser actions, constrain the account and environment, add an explicit transaction gate, and independently verify the outcome.

How does this relate to the rest of the DevDay release?

The DevDay 2026 hub maps the whole release. The companion Codex Cloud and Security Cloud analysis covers engineering, while RohitAI’s earlier Agents API guide explains the hosted harness.

Conclusion: autonomy should be budgeted, not granted

The Decisions API and computer-using Agents API point in opposite directions only if autonomy is treated as one switch.

In a good system, autonomy is a budget. The decision layer gets freedom to choose among known routes. The planner gets freedom to propose a method. The browser gets access to approved origins. Typed tools get narrowly scoped credentials. A human gets the final word on exceptional risk. The verifier and application ledger decide what counts as complete.

That design is less magical than a browser agent wandering the web until it declares success. It is also much closer to something a company can debug, audit, insure, and trust.

OpenAI supplied more of the stack at DevDay 2026. Builders still have to supply the boundaries.