Article

OpenAI Decisions API Beta: When to Use It Instead of Responses

Choose between OpenAI Decisions and Responses for support routing and classification, with request constraints, fallback costs and evaluation guidance.

Editorial illustration for OpenAI Decisions API Beta: When to Use It Instead of Responses: a document represents the research briefing. Not documentary evidence.

OpenAI released the Decisions API beta on October 6, 2026, according to its API changelog. For builders using generation calls to classify content, route support tickets or prioritize queues, it offers a dedicated fixed-answer interface. The current guide identifies it as public beta, with gpt-6-luna as the only supported model.

This expands access to an existing offering: OpenAI’s September 29 DevDay recap introduced Decisions in limited preview. The October 6 development is the beta rollout, not Luna’s model launch.

Choose by what the next step needs

Our recommendation is to evaluate Decisions for one classification step whose evidence and possible outcomes are already defined. Keep the rest of the workflow intact until the replacement meets the same error budget.

Required result

Candidate interface

Example

A condition’s probability, one category or a rubric score

Decisions

Assign a ticket to a queue.

Generated fields or an explanation in your own JSON schema

Responses with Structured Outputs

Extract order details and draft a reply.

A requested function and its arguments

Responses function calling

Propose an account lookup with parameters.

For support routing, a choice can assign a ticket to a queue; your code still decides whether to act or send it for review. A label is not a tool invocation or permission to issue a refund. OpenAI’s function-calling flow leaves tool execution to the application.

The Decisions guide defines predicate as a probability, choice as a selection from supplied options, and score as the probability-weighted mean of zero-based rubric levels. For overlapping labels, use separate predicates; for mutually exclusive destinations, use a choice.

Migration changes the request and error handling

Use POST /v1/decisions with model, input and questions. The API reference accepts text strings or user messages with text and inline base64 images. Remote image URLs, file IDs, assistant messages, tool calls and tool outputs are unsupported. It caps image parts at 128 per request.

That means a Responses conversation cannot simply be forwarded unchanged. Build a task-specific evidence payload, preserving the facts needed for classification. Include image retrieval and encoding in any end-to-end latency comparison.

Illustrative request body for a support queue, based on the documented schema; it has not been executed:

{
  "model": "gpt-6-luna",
  "input": "I need to change the delivery address for my order.",
  "questions": [
    {
      "type": "choice",
      "name": "queue",
      "instructions": "Select the team responsible for the requested change.",
      "choices": [
        {
          "value": "delivery",
          "description": "Shipping destinations and delivery arrangements."
        },
        {
          "value": "account",
          "description": "Account profile and sign-in changes."
        },
        {
          "value": "review",
          "description": "Ambiguous requests or work outside these teams."
        }
      ]
    }
  ]
}

Answers preserve question order, and names identify them. An individual answer can instead have type: "refusal", without a probability or choice. Handle that documented variant before reading answer fields; a refusal is not a negative predicate. Give refusals, timeouts and ambiguous classifications explicit fallback paths.

Independent questions can share a request; questions depending on earlier answers need separate requests, according to the guide.

Calculate savings after fallbacks

The Decisions rate is US$0.10 per million input tokens, with no output-token, cache-read or cache-write charges. Regional premiums and long-context input multipliers still apply. Standard Luna generation pricing lists $0.10 per million uncached input tokens and $0.50 per million output tokens.

Illustrative calculation, not a measured bill: assume one million text requests, each billed for 500 input tokens on either endpoint. Assume Responses additionally bills 20 output tokens per request, including any reasoning tokens. Exclude caching, regional and long-context multipliers, tools and retries.

Design

Calculation for one million original items

Token cost

Decisions only

500 million input tokens × $0.10 / million

$50

Responses only

$50 input + 20 million output tokens × $0.50 / million

$60

Decisions, then Responses for 20% of items

$50 + 0.20 × $60

$62

On those assumptions, removing output billing saves $10, or 16.7%, before fallbacks. If every item first uses Decisions and fraction f then incurs the full original Responses call, cost becomes $50 + f × $60. It beats the $60 baseline only when f is below one-sixth, approximately 16.7%.

This break-even point belongs to this example, not the product. Different prompts, billable outputs, cache usage or fallback models change it; integration and human-review costs are excluded. Multi-question and image workloads need actual usage accounting, not an assumption that extra questions are free.

OpenAI advertises roughly 10× faster answers than Responses in its release materials. The reviewed guide does not supply a reproducible workload or p95 latency distribution. Treat that as a vendor claim, not an end-to-end guarantee or a proportional cost saving.

Set thresholds against errors you can afford

A probability and an action threshold answer different questions. Scikit-learn’s threshold guidance explains why the threshold should reflect the costs of false positives and false negatives. A routing system that can tolerate a wrong queue may need a different policy from one escalating urgent service failures.

A proposed evaluation for the replacement step:

  1. Freeze the label definitions, rubric and input preparation. Tune thresholds on validation examples, then assess a separate labeled holdout.

  2. Compare wrong routes, missed urgent cases and review volume by category against the current system. Include out-of-scope requests and refusals, not only easy examples.

  3. Check whether predicted probabilities match observed frequencies using calibration curves. Do not read the separate confidence field as a demonstrated percentage accuracy.

  4. Run in shadow mode without changing live actions. Record billed usage, fallback frequency and median/p95 total latency under representative load.

Check the production project’s data controls

OpenAI’s endpoint-specific data policy lists up to 30 days of default abuse-monitoring retention. Eligible customers can use Zero Data Retention, with limitations: encrypted GPU-local cache tensors can persist for up to 24 hours, and image safety-review exceptions apply.

HIPAA use requires an executed Business Associate and Healthcare Addendum and the applicable account configuration. Regional processing is listed for the US and Europe (EEA plus Switzerland); regional storage is a separate capability. The data-controls table marks Private Retention with PSP and Safety Retention eligibility as pending confirmation. Do not infer those modes from ZDR support.

For the agreement and project-configuration checks, see our OpenAI HIPAA BAA setup guide. Verify Decisions access, quotas and processing-tier support for the actual project; the reviewed sources do not establish that every ordinary Luna limit or mode carries over.

Liquid d1 is a separate integration

Teams comparing providers can use our Liquid d1 guide. Liquid’s API contract uses state, keyed questions and Noul/Choice/Score types, and charges supplied images again for each question. It is not a drop-in request format for OpenAI. Compare accepted decisions at the same error budget and measured usage, rather than treating token rates alone as a task-cost benchmark.

Decision: pilot Decisions where a fixed answer is the whole deliverable. Retain Responses where extraction, explanation or generated tool arguments are essential. Promote the pilot only when its error rate, fallback cost and total latency justify the request-format change.

Methodology: AI-assisted analysis of primary documentation and release announcements, with availability, pricing and key links rechecked on October 6, 2026. No authenticated API calls, account-access checks or performance tests were conducted. Cost figures are illustrative calculations; evaluation steps are proposals.