On October 9, 2026, OpenAI published a case study describing Sophos’s use of Daybreak agents in managed detection and response (MDR). For security operations leaders, it offers a workflow to examine before deciding which investigations and response actions to automate.
Today’s publication is not the first disclosure of the headline results. Sophos’s May 28 release already reported 89-second automated response and 52% autonomous case closure. Those remain company-reported figures, not independently replicated findings.
According to OpenAI’s account, an agent assembles case evidence and threat intelligence; a planning model directs an investigation loop and produces recommendations for analyst review. Other agents can perform response work, while cases outside the agents’ remit go to people. The practical decision is where customer permission and analyst escalation sit between an investigation and a response action.
Read the two metrics separately
The October case study compares an earlier process averaging about 38 minutes with 89 seconds for cases using agents. Calculation: 38 × 60 = 2,280 seconds; (2,280 − 89) ÷ 2,280 ≈ 96%. This checks the rounded arithmetic, not the quality of the comparison.
The May release defines the 89-second interval as case creation to automated response for cases AI is authorized to resolve. Its 52% figure measures end-to-end case closure. These answer different questions: how quickly an eligible case receives a response, and how much of the workload is automated.
The October account does not establish the sample size, case mix, tail latency or equivalence of the comparison groups. It also does not identify the deployed model versions. The published case study therefore cannot isolate a model-specific causal effect.
For a buyer, the implication is to compare like-for-like case classes and report escalated cases separately. Do not multiply the automation share by the time reduction to estimate labor savings. Neither a fast response nor a closed case, by itself, establishes fewer missed threats or less customer disruption.
Customer authorization is not an agent permissions list
Sophos’s current response-mode documentation defines the service’s authority more precisely than the case-study shorthand. The buyer checks below are proposed evaluation questions.
Mode | Documented service authority | Buyer check |
|---|---|---|
Notify Only | Notifications and limited investigation; no containment or neutralization. | Who owns containment, including outside business hours? |
Collaborate | Investigation and recommendations; response requires written customer authorization. | Which actions need consent, and is the unreachable-contact fallback enabled? |
Authorize | Proactive containment with notification; full neutralization for MDR Plus. | Which actions still need internal analyst review, and how is recovery handled? |
Collaborate has an important optional fallback: if attempts to phone all customer-defined contacts receive no acknowledgment, Sophos may operate in Authorize mode when that option is selected. A mode name alone does not establish that the customer will approve every response in advance.
The separate MDR operations action catalog includes device isolation, account disabling and file deletion. Device commands require Live Response to be enabled; Microsoft 365 response actions require that integration. These are service capabilities, not proof that AI independently performs every listed action.
That distinction creates two reviews: what the customer permits the service to do, and what the service permits an agent to do without an analyst. Ask for an action-specific matrix covering permissions, review, notification and recovery. For a custom implementation, enforce those limits in the tools or policy layer, not only in the prompt.
A pilot should exercise the exceptions
Start with one defined case class in a non-production or recommendation-only evaluation. The following are proposed acceptance requirements, not tests performed by RohitAI or disclosed Sophos results:
Eligibility: Specify the evidence required for automation and what forces escalation. Include incomplete or conflicting evidence, not just well-contextualized cases. Record the automation-eligible share separately from the successful-closure share.
Authority: Exercise the selected response mode, denied actions and an unreachable contact. Record which permission authorized each action and whether human review occurred at the required point. Keep potentially destructive changes outside the initial agent scope.
Outcome: Separate case creation, first response, containment and closure timestamps. Have a reviewer assess closure correctness, reopened cases, missed threats, unauthorized actions and service disruption. Compare average and tail latency across matched case types.
Agree pass/fail criteria with the service owner before expanding authority. A faster average should not compensate for a required approval being skipped. RohitAI’s Ironclad workflow-evaluation analysis develops the same distinction between task scores and accepted operational outcomes.
Price an accepted case, not a single model call
For teams building their own integration, OpenAI’s current Daybreak API guide requires organization approval and project access to both the program and model. A request parameter does not grant access. Nor does model access grant permission to act on a customer’s systems. The earlier Daybreak access analysis provides background; use current documentation for implementation.
As checked on October 9, OpenAI lists GPT-5.6 Cyber at US$12.50 per million input tokens and US$75 per million output tokens. An illustrative request with 10,000 uncached input tokens and 2,000 billed output tokens costs (0.01 × $12.50) + (0.002 × $75) = $0.275. These are hypothetical inputs, not Sophos workload measurements or MDR service prices.
A case can need multiple calls and retries. For a pilot, divide all model, tool, infrastructure and review/rework costs across all attempted cases by independently accepted resolutions. Report critical-control failures separately; if no cases pass, report spend without a per-success ratio. Obtain service pricing and deployment-specific data-handling terms separately.
Sophos’s disclosed workflow gives buyers concrete questions to ask: which cases qualify, who authorizes each action, and what evidence makes a resolution acceptable. Those answers determine how much response authority to delegate.
Methodology: AI-assisted reporting and analysis of published primary sources checked October 9, 2026. Performance figures are company-reported; calculations use stated inputs. No hands-on evaluation or customer outcome was measured.
