Article

When Claude’s new workflow beta makes sense for document review

Claude Managed Agents can coordinate large document jobs. How to choose workflows or batching, check review coverage and compare costs per accepted document.

Editorial illustration for When Claude’s new workflow beta makes sense for document review: a controlled task flows from input to output. Not documentary evidence.

Anthropic introduced dynamic workflows for Claude Managed Agents on October 9, 2026, with large document reviews among the intended uses. For operations and research teams building review tools, the change lets Claude write a background program that assigns work to multiple AI agents and combines their results. The release is in beta.

An agent is AI software given a task and tools; a workflow coordinates those tasks. It can divide a document collection among agents, then bring their findings together. This is a developer service, not a new Claude chat feature: Managed Agents access is enabled by default for API accounts, with a beta header required on requests.

The decision is whether a job needs that coordination. Hundreds of files containing the same fields may suit simpler batch processing. A collection that needs separate investigations followed by a shared conclusion is a stronger candidate for workflows. Neither size nor agent count establishes that the result will be faster, cheaper or more accurate.

Choose by the work, not the number of files

The following is a practical decision framework, not a performance ranking. Anthropic’s earlier architecture guidance recommends starting with one agent and adding others when separable tasks or specialist contexts justify the coordination.

Your task

Starting point

Main tradeoff

Review a few documents, or reason about tightly connected material.

One agent.

Keep the reasoning together; establish a baseline before adding coordination.

Extract the same fields from independent files, without needing immediate answers.

Message Batches.

Requests run independently. Your application still needs to collect and check the results.

Ask specialists for findings, then question them further.

Direct subagents.

The coordinating agent can send follow-up messages.

Divide a large investigation into pieces, then reconcile the findings.

Dynamic workflows.

A program coordinates the agents; workflow threads do not accept follow-up messages.

For example, a proposed report-review workflow could examine each report for evidence, compare the findings and flag contradictions for a reviewer. That is a design example, not a tested integration. The useful output would include traceable evidence and unresolved questions, not just a combined summary.

The separate Message Batches API offers 50% lower input and output token rates than standard requests, but unfinished processing can expire after 24 hours. It is not a drop-in replacement for a coordinated workflow, and Managed Agents does not receive that batch discount.

Check which documents actually passed review

Anthropic’s workflow documentation says a run can finish with a completed status even when some work failed or a thread could not be created. Completion tells you the program ended; it does not establish that every document was reviewed successfully.

A practical acceptance record should distinguish three things: a supported finding, a supported finding that nothing relevant was present, and an unresolved document. An unreadable file must not quietly become a negative result. For each document, require:

  • A unique identifier and version, so retries do not count as extra documents.

  • The requested findings with page or section references that a reviewer can check.

  • An accepted or unresolved status, with the reason for any gap.

Check the collection-level conclusion separately. Accurate summaries of individual files do not, by themselves, prove that a comparison caught every contradiction. Agree in advance who accepts the result and how negative findings will be checked.

The current documented limits include 64 simultaneously working threads per run, a number the API does not guarantee, and 1,000 workflow-started agents over a run’s lifetime. That is not 1,000 concurrent reviewers or a promised completion time.

Count cost per accepted document

At USD list prices checked on October 10, Managed Agents charges for model tokens plus $0.08 per running session-hour. Standard Claude Opus 5.5 costs $4 per million uncached input tokens and $20 per million output tokens.

Illustrative calculation, not a test or quote: assume one standard Opus 5.5 session uses one million uncached input tokens and 100,000 output tokens across all planning, review and retries, plus one billable hour. Assume no searches, caching or pricing modifiers. The calculation is $4 + $2 + $0.08 = $6.08.

If 100 unique documents meet predefined acceptance checks, that is $0.0608 each. If only 80 pass for the same spend, it becomes $0.076 each—25% more—with 20 documents still unresolved. These assumptions do not estimate typical document-review usage.

Human review, document extraction, storage and external tool costs are excluded from that example. For a purchasing comparison, add them to all attempt charges and divide by unique accepted documents. Report the coverage rate alongside the cost: cheap accepted results do not excuse leaving difficult files unfinished.

For usage accounting, use the session’s authoritative list-cost total. Individual thread costs are rounded separately and omit session runtime. Overlapping thread activity counts once for billable session time, so do not multiply the hourly rate by the number of agents.

Set the budget and stop behavior before a pilot

A session budget covers all its threads and workflows at public list rates. Set it when creating the session. Enforcement blocks new model requests once the cap is reached, but already admitted requests finish; overshoot can include one request per working thread. Leave spending headroom rather than treating the cap as an exact final bill.

Interrupting the main conversation does not end its workflows. Runs may pause or continue. The documentation directs clients to ask the agent to stop its runs, though budget-paused sessions reject new messages until the budget is raised or removed. Either change can also resume work. Require a verified stop procedure rather than assuming the chat interrupt is a stop-all button.

Check document eligibility first. Managed Agents stores session history, state and outputs and is not currently eligible for Zero Data Retention. For applications fetching documents online, our Managed Agents web-access guide explains the separate question of restricting which sources they can reach.

Start a pilot with authorized, representative documents and fixed acceptance criteria. Compare a workflow with one agent—or batching where appropriate—on the same inputs. Record omissions, unsupported findings, elapsed time, all costs and reviewer effort. Scale only if the resulting evidence justifies the added coordination; the reviewed sources establish this beta’s capability, not its return on document-review work.

Methodology: AI-assisted reporting and analysis based on Anthropic’s published release notes and documentation, rechecked October 10, 2026. No hands-on benchmark was run. The decision framework and cost calculation are illustrative.