Article

Cloudflare Clef on Workers AI: Routing Requests, Probabilities and Costs

How to use Clef and Clef-flash for typed agent routing: Workers AI request formats, review thresholds, hosted limits and input-cost calculations.

Editorial illustration for Cloudflare Clef on Workers AI: Routing Requests, Probabilities and Costs: a controlled task flows from input to output. Not documentary evidence.

Cloudflare announced Clef and Clef-flash on October 1, 2026, releasing decision models on Workers AI and downloadable weights under Apache-2.0. For builders of support routers, classifiers and agent workflows, the change is a purpose-built interface for scoring known choices instead of asking a general language model to write an answer.

The Clef model card describes a single forward pass that assigns probabilities to the supplied options, without generating free-form prose. A support router can send a ticket to billing, technical support or human review. Your application still decides which predictions to accept and which actions are permitted.

Use it when the choices are known but the meaning is not

Keep arithmetic, account-status checks and explicit business rules in ordinary code. Consider Clef where the input is unstructured but the possible outcomes are bounded: which queue should receive this ticket, or does this message describe an outage? Keep a general LLM for drafting replies and open-ended planning.

In a retrieval workflow, first gather the relevant evidence, then score the allowed routes, then apply application policy. The Cloudflare AI Search guide covers the upstream retrieval stage; it does not establish an automatic Clef integration or included Clef inference charges.

Clef versus Clef-flash: price is only the first filter

The hosted Clef documentation and Clef-flash documentation list the following specifications and USD input rates, checked October 1. Backbone names come from the pinned Clef and Clef-flash model cards.

Property

Clef

Clef-flash

Parameters / backbone

27B / Qwen3.8-27B

9B / Qwen3.5-9B

Hosted context window

65,536 tokens

65,536 tokens

Price per million input tokens

$0.24

$0.09

Workers AI identifier

@cf/cloudflare/clef

@cf/cloudflare/clef-flash

Cloudflare’s launch evaluation table reports 97.43 versus 66.77 macro-F1 for Clef and Flash on CLINC150+OOS, an intent-classification task including out-of-scope inputs. Flash edges ahead on its typed BFCL function-selection evaluation, 98.76 versus 98.47 case-exact accuracy. These are vendor results, not independent tests or full agent-success rates. The practical inference: evaluate the task you need, especially unfamiliar inputs, rather than picking Flash solely for price.

A minimal Workers AI routing request

The hosted input schema requires model, state and questions, plus instructions for each question. It permits 1–64 questions. Use choice for 2–255 named options, noul for a yes/no probability, and score for 2–10 ordered levels.

This illustrative function adapts the documented Workers AI binding. It assumes the Worker already has an AI binding named AI and receives bounded ticket text. It has not been executed; it returns a classification response and performs no routing action.

async function classifyTicket(env, ticketText) {
  return env.AI.run("@cf/cloudflare/clef-flash", {
    model: "clef-flash",
    state: ticketText,
    questions: {
      team: {
        type: "choice",
        instructions:
          "Choose a queue. Use review for ambiguity or missing evidence.",
        criteria: {
          billing: "Invoice or payment enquiries",
          technical: "Service faults or configuration help",
          review: "Unclear, unsupported or insufficient information"
        }
      },
      urgent: {
        type: "noul",
        instructions: "Does the evidence describe an urgent service interruption?"
      }
    }
  });
}

To evaluate the larger model, change both the binding identifier to @cf/cloudflare/clef and the request’s model value to clef. Keep the same labels and evidence when comparing results.

Validate the response before accepting a route

The output schema returns answers keyed by question ID. Read response.answers.team.choice and response.answers.team.probabilities. Urgency is the number at response.answers.urgent.noul: testing the whole urgent object as a boolean would not test urgency.

  • Validate the expected question IDs, answer types and option names. Reject missing, non-finite or out-of-range probabilities; check that distributions sum to one within a justified rounding tolerance and that the chosen option has maximal probability.

  • Treat review as a real outcome. Also defer when evidence is incomplete, class-specific acceptance thresholds are not met, or the request times out, is throttled or returns malformed data. A failed request must not default to a business action.

  • Set thresholds using held-out examples and the cost of each error. Compare accepted-case errors against the fraction sent for review; a router that avoids mistakes by deferring everything is not a useful automation result.

A probability of 0.90 is not, by itself, evidence of 90% accuracy on your tickets. Calibration research explains why confidence and empirical correctness must be compared. It does not establish whether Clef is calibrated for a new workload.

If you add a severity score, inspect its distribution as well as its average. The reference adapter computes an expected level. In a mathematical illustration, half the probability at level 0 and half at level 3 gives the same mean, 1.5, as half at level 1 and half at level 2. Only the first assigns any probability to level 3. Neither distribution is reported model output.

For audit records, retain a privacy-appropriate evidence reference, model identifier, schema and policy versions, probabilities, acceptance decision and action receipt. Typed answers are easier to record, but do not supply a reasoning trace or authorization. The bounded-autonomy guide develops that architectural separation.

Do not copy the local contract into a hosted request

The open-weight release and Workers AI share a decision format, but their documented interfaces differ:

  • Instructions: Workers AI requires them; the local helper can fall back to the question ID.

  • Context: hosted pages list 65,536 tokens, while the local encoder defaults to 16,384. Budget for questions and other overhead, not just ticket text. Bound the evidence yourself rather than relying on truncation.

  • Media: the hosted schema accepts up to four embedded PNG, JPEG or WebP images, not remote image URLs. Limits are 4 MiB and 16 megapixels per image, 8 MiB decoded total and 13 MiB for the request. The local model card supports video frame arrays; hosted video acceptance is not established by its raw schema.

What a Flash-first cascade would cost

Illustrative arithmetic, not a measured bill: assume one million decisions, each consuming 1,000 billable input tokens including question/schema overhead, with identical token counts for both models and no retries. That is one billion input tokens. Using the published Workers AI rates:

Strategy

Input-inference cost

Clef for every decision

$240

Flash for every decision

$90

Flash for all; Clef again for 20%

$138

If q is the fraction re-evaluated by Clef, cascade cost is $90 + $240q. It is below direct Clef only when q is less than 62.5%; at 20%, the reduction is 42.5%. This compares input spend, not equal-quality outcomes. Both models may fail on the same difficult cases, and serial fallback adds latency.

The illustration excludes the shared free allocation, plan and execution costs, retrieval, media-token differences, downstream generation, human review, discounts and taxes. It is not a claim of all-in savings.

Start with a labeled routing evaluation

Before connecting predictions to actions, compare both models with your current rules or classifier on a frozen set of labeled tickets. Keep calibration and final evaluation data separate. Include rare queues, unfamiliar requests, relevant languages and incomplete evidence. Measure per-class precision and recall, costly mistakes, deferral rate, and end-to-end p50/p95 latency at your expected concurrency and input sizes. Test the fallback path separately.

Confirm effective account quotas and access before sizing production traffic; this guide has not verified them through inference calls. If evaluation reveals a domain-specific gap, Cloudflare currently offers hands-on fine-tuning with its forward-deployed engineers. Self-service is planned, not documented as available, and no launch date or fine-tuning price was verified.

Methodology: AI-assisted analysis of Cloudflare’s announcement, hosted schemas, pricing and pinned model-card/reference-code revisions, checked October 1, 2026. No inference calls or workload benchmarks were run. The code is an untested documentation adaptation; cost figures are explicit calculations, not observed bills.