Anthropic’s Building effective agents, published on December 19, 2024, gives builders and operations teams a useful starting point: use the simplest system that meets the task. Its page now cautions that the tooling has changed. This is an evergreen decision guide using current documentation checked on October 5, 2026—not a new product announcement.
If a person already has the facts and needs wording, start with a draft in chat. If software can prescribe the necessary lookups, use a workflow. Consider an agent when choosing what to investigate next genuinely requires adaptation—and when you can check the result.
Choose who controls the next step
A chatbot describes a conversational interface, not the machinery behind it. That interface can front a fixed workflow or an agent. Here, “chat-only” means a model drafts from facts a person supplies, without retrieving records or taking external actions itself.
Anthropic’s architectural distinction concerns control: workflows orchestrate models and tools through predefined code paths; agents let the model direct the process and tool use. Tool access alone does not establish open-ended autonomy.
A workflow need not be a straight line. Microsoft’s architecture documentation includes branching, parallelism and error handling under code-directed control. An LLM can classify a case or draft within a step; predictable control flow does not mean identical wording or guaranteed correctness.
One delayed delivery, three designs
The same delayed-delivery reply can be drafted from human-supplied facts, a fixed lookup path or model-directed investigation.
This is an illustrative design comparison, not an executed test. Assume an authenticated support operator may access the relevant customer’s order, official carrier records and applicable policy. The deliverable is a reply draft with supporting references for the reviewer. Sending, refunds and new delivery commitments are outside the task.
Apply the same acceptance rule to each design: material claims must match authorized records and their timestamps; policy must apply to this order; conflicts and missing facts must remain explicit. A polished answer without that evidence does not pass.
Design | Who gathers evidence and chooses steps? | What happens when evidence is missing? | Good starting fit |
|---|---|---|---|
Chat-only draft | The operator retrieves order, carrier and policy facts. The model writes from that packet. | The draft flags the gap; the operator investigates or asks for clarification. | Occasional cases where human retrieval is manageable. |
Fixed workflow | Code runs scoped lookups, freshness checks, drafting and validation along declared branches. | Known timeouts get bounded retries; missing identifiers or unresolved conflicts go to a person. | Repeated cases with lookups and exceptions you can specify. |
Model-directed agent | The model selects among narrow read tools, using results to choose further investigation. | It may inspect tracking history or request context, then stops or escalates within defined limits. | Cases where useful next queries vary and the evidence can still be checked. |
All three end at human review. The distinction is who chooses the evidence-gathering path, not whether the system is allowed to send the reply.
For example, suppose an order page says dispatched while a newer carrier record reports an exception. A workflow can follow a predefined reconciliation rule or escalate. An agent could inspect the event history through an authorized read tool. Neither design should turn an unresolved conflict into a promise of arrival tomorrow.
An exception is not automatically a reason to add an agent. A practical hybrid to evaluate is a fixed workflow for ordinary cases, with only unresolved investigation passed to a bounded read-tool agent. Return its evidence to the same checks and reviewer. LangChain’s workflow and agent patterns provide implementation context; this hybrid is a proposed design, not a demonstrated cost or accuracy improvement.
Give investigation autonomy without spending or sending authority
Claude’s client-tool interface separates a model’s tool request from execution: application code runs the function and returns a result. For this example, make that code enforce which customer and order may be queried. Do not let a model-supplied identifier expand access.
Expose order lookup, tracking history and policy retrieval—not a general shell, unrestricted browser, send button or refund endpoint. Keep backend credentials scoped to those operations. Read-only is not harmless: retrieving another customer’s records is still a disclosure.
If using Claude Agent SDK, its current permission documentation has two important qualifications. allowedTools / allowed_tools pre-approves listed tools; it does not remove all unlisted tools. And canUseTool is not a universal gate: calls approved earlier in evaluation skip it. Inspect the effective tool surface, permission mode and deny/ask rules; use a PreToolUse hook for checks that must run before those decisions. Confirm behavior against the deployed version.
Approval is also not isolation. Anthropic’s secure-deployment guide distinguishes permission gates from sandboxing. Restrict filesystem access and network destinations separately, and keep credentials outside the model’s context. If sending is automated later, bind approval to the exact recipient, message and action parameters; changed content needs a new decision.
Retrieved emails, pages and API text are evidence, not instructions. Anthropic’s tool-result guidance treats such content as untrusted and recommends keeping it in tool_result blocks rather than promoting it into system prompts or user instructions. An email demanding a refund cannot authorize one. This handling reduces exposure; it does not guarantee immunity to prompt injection.
Set stopping and recovery rules before the first run
Write a short operating contract for the delivery task:
Limits: Choose maximum tool calls, elapsed time and estimated spend from the service’s requirements. Enforce them in the application, including retry paths, and define the handoff when any limit is reached.
Escalation: Stop for missing authorization, unresolved policy conflicts, repeated failures without new evidence or exhausted limits. Assign a person or queue to receive the evidence collected so far.
Audit record: Retain source identifiers, timestamps, tool requests/results, validation outcomes and the reviewer’s decision in access-controlled logs. Minimize customer data, exclude secrets and record success, escalation and limit termination separately.
The Agent SDK’s loop controls include maxTurns / max_turns and maxBudgetUsd / max_budget_usd; both are documented as unlimited unless configured. Turns count tool-use round trips, not individual calls or elapsed time. The separate API task_budget feature is advisory; max_tokens caps generated tokens per request. None is an all-in operating-cost ceiling.
If a later, separately authorized send times out, do not assume it failed and send again. Reconcile the durable action identifier with the provider’s receipt or status before replaying it. AWS’s idempotent-retry guidance explains how a service can recognize the same logical request. Verify the actual messaging service’s supported keys and retention window; if the outcome remains ambiguous, escalate. Stopping an agent does not undo an external effect.
Compare accepted replies, not fluent text
Before expanding access, evaluate permissioned or synthetic cases: ordinary delay, stale tracking, conflicting records, unavailable carrier, missing order identifier and hostile instructions embedded in retrieved text. This is a proposed evaluation, not reported test results.
Keep the acceptance rules identical across designs. Record evidence accuracy, unsupported promises, unauthorized-action attempts, escalation, elapsed time, staff effort and spend. Include the operator’s retrieval time in the chat-only baseline, and disclose any differences in information supplied. Otherwise the comparison hides work done by humans.
A useful accounting rule for that evaluation is: cost per accepted reply = total operating cost across all attempted cases ÷ replies that pass the agreed checks. Include failed and escalated cases in the cost total, and report their counts separately. If no replies pass, report that outcome and total spend rather than a ratio.
Count model requests, external-tool fees, infrastructure, retries and staff retrieval/review time, without double-counting labor. The SDK’s cost-tracking documentation calls total_cost_usd and costUSD estimates, not authoritative bills. Deduplicate shared message IDs and do not add cumulative resumed-session totals as though each were new spend. Reconcile model spend with provider billing and add costs outside the SDK.
For coding tasks, substitute an appropriate acceptance check for the support evidence review. RohitAI’s legacy-migration parity-harness guide shows how to define behavioral comparisons before accepting agent-produced code.
Start with the least complex candidate that passes those checks. Add model-directed investigation only when the simpler design leaves useful work unresolved and the evaluation justifies the added operating burden. If no design meets the evidence or access requirements, keep the task human-led.
Methodology: This AI-assisted guide draws on published Anthropic, Microsoft, LangChain and AWS material, checked on October 5, 2026. The delivery comparison, hybrid and evaluation method are proposed analysis. No carrier integration, hands-on benchmark or measured savings are claimed.
