Cloudflare announced AI Search general availability on October 1, 2026, giving builders of site search and retrieval-augmented assistants a new default configuration. New AI Search instances combine semantic vector retrieval and full-text matching by default. Cloudflare says usage-based AI Search billing begins November 1, 2026.
The service searches connected content—direct uploads, websites you own and R2 buckets—not the open web at large. Those source options make it a candidate for managed knowledge-base search. Before adopting it, choose the retrieval mode, check the corpus limits and estimate the costs beyond retrieval.
Choose hybrid for the query mix, not just the default
Hybrid retrieval itself is not new: Cloudflare introduced it in April. The October change makes it the default for new instances; it does not establish that existing instances were automatically migrated. Check an existing instance’s settings before planning a change.
Hybrid runs semantic and keyword retrieval in parallel, then combines their results; Reciprocal Rank Fusion is the default merge method. The following are evaluation starting points based on the documented search modes, not measured winners for your data.
Typical query need | Mode to evaluate first | What to check |
|---|---|---|
A named error code, product ID or function | Keyword | Does tokenization preserve the term that distinguishes the right document? |
A question phrased differently from the documentation | Vector | Does semantic matching retrieve the relevant passage without relying on shared wording? |
A specific identifier inside a broader troubleshooting question | Hybrid | Does the combined ranking retain the exact reference and useful explanatory context? |
Keyword matching is not a byte-for-byte lookup guarantee. Cloudflare offers stemming and trigram tokenizers; changing the tokenizer or index_method triggers a full reindex. Its keyword-index documentation also sets a Workers Paid limit of 500,000 files per keyword-enabled instance, versus 1 million for vector-only indexing.
Capacity implication: a 750,000-file corpus fits the documented vector-only file ceiling but exceeds the keyword-enabled ceiling for one Paid instance. Selecting vector-only requests on a hybrid index does not remove that index-level limit. Splitting the corpus or retaining vector-only indexing becomes an architectural decision, not merely a per-query preference.
Included embeddings do not include every model call
Workers AI embedding and reranking calls made by AI Search are bundled into AI Search pricing and no longer appear on the Workers AI bill or in AI Gateway logs. Generation, query rewriting and external-provider usage remain separately accounted for through the applicable service. That is the boundary in the GA announcement, not a promise of free inference.
Cloudflare’s search endpoint returns retrieved chunks; chat completions adds a generated response. Use search alone when an application needs evidence passages or a results list, and audit rewriting and provider settings before estimating the bill.
Reranking is off by default. It can reorder retrieved candidates, but the extra processing may add latency. Compare it enabled and disabled; bundling the Workers AI charge does not establish a quality or response-time benefit. Because the bundled operations disappear from Gateway logs, application-level retrieval timing remains useful.
Decide how images and scanned PDFs enter the index
Image input, native image embeddings and OCR solve different problems. The REST API accepts images in search and chat requests. A supported multimodal encoder embeds an image directly; a text-only encoder first turns it into a caption.
Cloudflare lists @cf/qwen/qwen3-vl-embedding-2b and google-ai-studio/gemini-embedding-2 as multimodal options. Its default embedding model is text-only. Image-query support therefore does not mean a new instance automatically uses native visual embeddings.
OCR extracts text from scanned documents. It is available on every account but remains disabled until enabled with indexing_options.use_ocr; changing that setting triggers a full reindex. The file-size limits are 10 MiB for plain-text/code files and OCR-enabled PDFs, but 4 MiB for PDFs without OCR and other converted formats. Oversized files are skipped and logged as errors.
The embedding model is selected when creating an instance, whereas the generation model can change later. For a switch from caption-based retrieval to native image embeddings, plan a separate-instance comparison before cutover. Check whether representative scans, diagrams and tables retain the information needed to answer real questions; format support alone does not demonstrate extraction fidelity.
For the broader encoder-migration decision, see RohitAI’s Cohere Embed 5 routing and reindexing guide. That is related architecture guidance, not a claim that Cohere’s encoder is supported by AI Search.
Estimate the November bill at account level
Illustrative monthly calculation using the current USD rate card: one text-only account, otherwise-unused allowances, and the metered usage below. Allowances are per account, not per instance; the two query allowances are separate.
Meter | Assumed usage | Included monthly | Overage calculation |
|---|---|---|---|
Ingestion | 20 million tokens | 5 million tokens | 15 × $0.75 = $11.25 |
Storage | 25 GB-month | 10 GB-month | 15 × $2 = $30.00 |
Semantic/vector/hybrid | 100,000 queries | 1,000 queries | 99 × $0.75 = $74.25 |
Full-text | 100,000 queries | 1,000 queries | 99 × $0.10 = $9.90 |
Retrieval-service subtotal: $125.40. Rates apply per million ingestion tokens, per GB-month and per 1,000 queries. This is arithmetic, not an observed invoice. It excludes images/OCR, generation, rewriting, external providers, Workers plan/application charges, taxes and discounts.
Image processing adds $0.50 per million tokens; OCR text counts toward base ingestion and image processing, sharing the 5-million-token allowance. The pricing documentation does not fully illustrate mixed-corpus allowance allocation. Confirm that metering before forecasting image-heavy costs; there is no documented per-page price here.
Use final indexed chunks for ingestion estimates. Overlap repeats text and can increase billable tokens; larger retrieved context can also raise generation cost. Budget refreshes and planned reindexing, and use metered storage rather than treating original file bytes as the storage bill.
What to validate before switching traffic
Label representative queries. Include identifiers, paraphrases, mixed queries, scans, image-dependent questions and cases where the corpus has no answer.
Compare retrieval paths. Hold the corpus, chunking and filters constant. Compare compatible modes and reranking settings; track whether relevant evidence reaches the top results, answer support where generation is used, and median/tail latency.
Make cutover reversible. Record index and encoder settings, inspect ingestion errors and preserve the incumbent path until the replacement meets your acceptance criteria.
Forecast from service usage. Do not assume one user action always equals one metered query. The reviewed rate card does not fully establish fan-out, cache-hit or retry accounting.
Hybrid is a useful evaluation baseline for a knowledge base that mixes exact references with natural-language questions. Retain a simpler retrieval mode where it meets your acceptance criteria, and choose OCR or native image embeddings because the source material requires them—not merely because GA makes them available.
Methodology: This AI-assisted decision guide uses Cloudflare’s published announcement and documentation checked October 1, 2026. The announcement establishes a calendar date, not an exact release time. The cost example is calculated from stated assumptions. No hands-on API tests, retrieval benchmarks or customer-outcome measurements were performed.
