On October 6, 2026, OpenAI reported its Ironclad research collaboration on computer-use agents for contracting. For legal-operations leaders and enterprise AI teams, the practical question is what evidence would justify a workflow-configuration pilot.
A contracting-agent pilot should test approval routing and reapproval before a person authorizes publication of the configuration. Evaluate the reusable process the agent creates, not just whether it finishes an editing session.
What the research scores measure
OpenAI describes 11 tasks with 8–50 grading criteria each, using synthetic training tasks and Ironclad-hosted environments. Its reported comparison is:
Reported metric | GPT-5.6 Sol · High reasoning | GPT-6 Astra · Max reasoning |
|---|---|---|
Mean rubric score | 41.6% | 55.0% |
Simulated average attempt time | 37.0 minutes | 19.2 minutes |
These are partial-credit scores, not percentages of contracts automated. Reasoning settings differ; this is not a matched-compute comparison. The timing footnote describes simulations based on assumed processing and generation speeds, not measured customer savings.
Calculation from those values: 55.0 − 41.6 = 13.4 percentage points; 13.4 ÷ 41.6 ≈ 32.2% relative improvement. Neither quantity measures savings per accepted workflow.
The report does not supply a complete reproducible task-and-rubric package or uncertainty intervals. OpenAI says training and evaluation excluded nonpublic Ironclad customer data, OpenAI customer data and OpenAI’s internal contracts.
Acceptance must cover the rules after the edit
Ironclad’s workflow configuration documentation explains that later workflows follow the published configuration. That makes the generated configuration the useful unit of acceptance: a routing error could recur in subsequent requests. This is a design risk to test, not an incident reported in the research.
Reapproval is a concrete example. Ironclad’s approver documentation distinguishes when approval is initially required from when it resets. It supports resets after document or property changes; selecting no reset condition prevents resets. The documentation specifically describes Finance reapproval when the contract-value property changes.
A pilot should therefore exercise both the first approval and a later edit. The following is a proposed acceptance matrix, not a reconstruction of OpenAI’s unpublished rubric. The legal-operations owner must define the expected outcomes from the organization’s policy.
Case | What to exercise | Evidence to retain |
|---|---|---|
Spending boundary | Requests below, exactly at and above the approval threshold. | Stored condition and resulting approver assignments for each case. |
Edit after approval | Change the contract-value property or document after approval. | Whether the expected reapproval occurs; compare with an unchanged control. |
Clause selection | Change the selected jurisdiction, then inspect the resulting clause. | Output compared with the owner-approved expected wording for each selection. |
Incomplete input | Omit a field needed to select a route. | The specified stop or clarification behavior, rather than a guessed value. |
Authority boundary | Encounter a denied action or document text requesting wider access. | No unauthorized action, with an explicit handoff when needed. |
Use partial credit to diagnose errors, but make required controls separate pass-or-fail checks. Correct field labels should not compensate for a missing required approval. This is a proposed pilot decision rule, not a claim about how OpenAI weighted its evaluation.
Research access and product availability are separate
GPT-6 Astra’s model documentation lists computer use through the Responses API, but no fine-tuning support. General model access is not a self-serve version of the Ironclad training collaboration.
OpenAI’s collaboration application asks software companies for a failing workflow, success criteria, domain expertise and a secure research environment; it also asks about containerized software instances. Applying does not establish participation.
Ironclad’s October release notes schedule Workflow Designer Agent availability for October 8, 2026. As of October 6, that is a planned release, with the company setting off by default and admin enablement required. The notes instruct users to review changes before saving or publishing.
The Workflow Designer Agent overview requires Workflow Designer permissions and excludes tasks such as template formatting, smart-tag placement and manually skipping runtime steps. These product documents do not establish that it uses the evaluated Astra setup; the research score should not be transferred to it.
What to require before a pilot goes live
Start with one permissioned draft configuration and a named owner for its business rules. In the initial pilot, keep live publication, signing and counterparty communications outside the agent’s authority. Let the owner review the stored configuration and the acceptance cases before authorizing deployment.
For a custom computer-use integration, OpenAI’s execution guidance calls for an isolated environment, restricted sites and actions, bounded runs and verification of actual outcomes. Screen and document text must not expand the agent’s permissions. Decide which contract fields and screenshots the evaluation may expose; the research’s data choices do not establish your deployment’s retention or confidentiality terms. The related agentic privacy analysis offers a data-flow review framework.
Keep the starting environment, model and reasoning setting, tool interface, rubric and run limits fixed when comparing candidates. Inspect both saved configuration state and behavior on held-out input cases. Repeat runs to expose variability and retain failures, not just successful demonstrations. OpenAI’s evaluation guidance distinguishes trace-based diagnosis from repeatable dataset evaluations; traces can explain how a wrong approval route arose.
The evidence packet should contain the original requirements, configuration changes, case-by-case outcomes, execution traces, actual elapsed time and reviewer corrections. For organizing versioned tasks, environments and graders, see the RL environments guide.
For pilot accounting, use total model, tool, environment, retry and review/rework costs across all attempts divided by accepted configurations. Report critical-control failures separately. If no configuration is accepted, report the spend without a per-acceptance ratio. This proposed measure keeps the review burden visible; it is not a cost result from the collaboration.
Advance only when the owner can explain why the generated configuration follows the intended policy across the agreed cases. A higher research score can motivate that evaluation; the deployment decision needs its own evidence.
Methodology: AI-assisted reporting and analysis of official sources checked October 6, 2026. Benchmark figures are vendor-reported; calculations use published inputs. No hands-on evaluation or customer outcome was measured.
