Article

Liquid d1 Adds Vision: When to Use a Decision Model Instead of Chat

A guide to Liquid d1’s paid vision API, per-question image costs, and the checks needed before replacing chat-based classifiers.

Editorial illustration for Liquid d1 Adds Vision: When to Use a Decision Model Instead of Chat: a geometric block represents a model release. Not documentary evidence.

Liquid AI announced vision support for d1 on October 5, 2026, extending its earlier text experiment with image-based decisions on the Liquid API. For builders using chat models to classify inputs, route work or inspect images, the change offers a hosted call that returns probabilities instead of composing an answer.

The useful migration target is one bounded decision, not an entire assistant. Consider d1 when the possible outcomes are known and the necessary evidence is already present. Keep a generative model when the task needs explanations, new content or complex reasoning. Liquid’s migration guide makes that task distinction; whether a replacement meets your error budget still needs evaluation.

Choose the decision step

The interface offers three primitives: Noul for a yes/no probability, Choice for a category, and Score for an ordered rating. Liquid’s guide maps these to classification, routing and scoring. The following is a proposed task-fit framework, not a report of tested integrations.

Your task

Candidate design

What to retain outside d1

Assign a request to one team

Choice with clearly separated categories and an out-of-scope route

Queue ownership and an explicit fallback for ambiguous cases

Apply several independent labels

A separate Noul question for each applicable condition

A multi-label policy; do not force overlapping labels into one Choice

Triage a visual defect or policy concern

Noul for a binary criterion; Score for a defined severity rubric

Review thresholds, representative labels and escalation capacity

Explain a defect, draft a reply or investigate missing evidence

Keep generation or investigation in a separate step

The model or workflow that can produce the required deliverable

A short final label does not make the underlying problem easy. Compare against the system you actually run, including deterministic rules or a trained classifier where appropriate—not only against an expensive chat prompt.

Deployment is also a deciding factor. Liquid’s model library lists d1 as API-only, without downloadable formats or a training path. If inputs must stay on your infrastructure, assess a different local model. Our guide to local typed-decision serving covers that separate option; it is not a way to self-host d1.

Vision access is endpoint-specific

The direct API reference specifies paid d1 for images; d1:free is text-only. Send state, named questions and images to POST https://api.liquid.ai/decisions/v1/systemone. Images require base64 data URLs or base64 objects, not remote URLs. The documented SDKs do not yet support images, so the examples use HTTP directly.

The same reference sets these request limits:

  • JPEG, PNG, WebP or GIF; up to eight images.

  • Entire JSON body below 4.5 MB, including base64 data.

  • At most 10,000 total 32×32 patches; aspect ratio no greater than 100:1.

Liquid’s launch announcement describes Vercel and OpenRouter as text-only routes, with vision planned but no dated rollout. OpenRouter’s current endpoint metadata still lists text input. Do not infer image support from the shared model name.

There is a discovery caveat: the public Liquid model catalog returned only the free text model when checked on October 6. That does not disprove the paid offering, but it is not an entitlement check. No authenticated inference was performed for this article. Before committing, confirm account access, direct-API context limits and quotas, data-handling terms and acceptable latency.

Budget images per question, not per request

Liquid’s launch calculation uses US$0.04 per million input tokens, also listed by Vercel. Billing excludes generated output tokens; a structured response still comes back.

Under the documented image accounting, every question is charged for its text and every supplied image. Image tokens are calculated as:

patches = ceil(width / 32) × ceil(height / 32)
image_tokens = ceil(1.5 × patches)

Illustrative calculation—not an observed bill: a 1024×1024 image yields 1,536 tokens per question. At the listed rate, the image-only charge for one million requests is:

Images per request

Questions per request

Calculated image-only cost

One 1024×1024 image

One

US$61.44

One 1024×1024 image

Three

US$184.32

Two 1024×1024 images

Three

US$368.64

Arithmetic: 1,000,000 requests × 1,536 tokens × image count × question count ÷ 1,000,000 × US$0.04. These examples assume unchanged image dimensions and exclude text, retries, other fees and escalation work.

The design implication is that bundling questions saves network round trips, not necessarily image charges. If only one question needs a picture, compare a separate image request with a combined request that sends the picture to every question. Include the extra round trip in that comparison.

Validate probabilities and the action policy separately

Liquid describes the outputs as calibrated, but that is not evidence of calibration on your labels. A reliability check compares predicted probabilities with observed outcomes across labeled cases. A field named confidence is not, by itself, a measured correctness rate.

Proposed evaluation before migration:

  1. Choose one existing decision and label representative cases. Include rare categories, ambiguous inputs and the image conditions you expect in production. Version the question wording, rubric, preprocessing and endpoint.

  2. Select thresholds on validation cases, then assess the frozen policy on a separate holdout. Threshold tuning should reflect error costs: missing a defect and unnecessarily sending a good part for review need not have equal consequences.

  3. For a defect-probability gate, use a lower boundary for passing a case, an upper boundary for flagging it, and a review region between them. Choose both boundaries from the evaluation and available review capacity; no universal cutoff is established here.

  4. Compare the whole workflow with your current baseline at similar error or review-volume constraints. Track missed cases, false alarms, escalation fraction, failed requests, billed inputs and p50/p95 latency—not just aggregate accuracy.

Score also needs careful interpretation: it is a probability-weighted position on a zero-based rubric, not necessarily the most likely level. Noul near 0.5 represents yes/no uncertainty, not medium severity. The response reference explains the distinction.

Keep authority in application code. A favorable probability must not override tenant boundaries, tool permissions or transaction rules; missing fields and timeouts need an error path. A decision model can choose among permitted branches without becoming the permission system. The broader workflow-versus-agent guide explains where that decision step fits.

What the launch benchmarks do—and do not—establish

Liquid reports 85–97% inspection accuracy across four VisA categories and 200–300 ms text decisions. Its methodology describes one run per application per model, list-price costing without prompt-cache discounts, and a good reference part shown alongside the inspected part. These are vendor results, not independently reproduced findings.

For an inspection pilot, preserve the reference context when comparing quality and budget whatever image packaging you actually send. The reported text timing is not a vision latency guarantee, and the selected comparisons do not establish a win over your optimized classifier.

Start with one queue or visual-triage gate in shadow mode, without changing live actions. Move it into service only if its errors, review load, end-to-end latency and total cost meet your requirements. If the decision needs evidence the input does not contain, improve that evidence path before changing the model.

Methodology: AI-assisted reporting and analysis based on published documentation and provider metadata checked October 6, 2026. No d1 inference or hands-on benchmark was run. Cost figures are arithmetic under stated assumptions; the evaluation plan is proposed, not executed.