On October 6, 2026, Mistral announced a public API preview of Mistral Large 4, giving coding-agent builders, document-application teams and enterprise model buyers a new hosted model to evaluate. The API preview is available now; Mistral says it plans to release weights by the end of October. That is a future release target, not present-day self-hosting availability.
Start a bounded evaluation if hosted processing is acceptable and you need to compare workload quality. If the purchase depends on local operation or license rights, wait for released artifacts and terms. A successful API pilot can inform that later decision without settling it.
API preview now, weights later
Your decision | Current evidence | Practical next step |
|---|---|---|
Compare hosted model quality | Use authorized sample tasks; confirm your account access and limits. | |
Run the model privately | Wait for a released checkpoint, its license and supported serving configuration. | |
Replace a production dependency | Keep a fallback and validate stability and support terms before migration. |
Mistral’s model documentation describes a multimodal mixture-of-experts model with 1.05 trillion total parameters, 49 billion active parameters, a 1.6 billion-parameter vision encoder and a 1M-token context window. Listed features include function calling, structured outputs and document Q&A. These are documented specifications, not independently measured capabilities.
Make the preview evaluation repeatable
Mistral’s lifecycle policy allows preview updates without notice, does not guarantee a move to general availability and specifies one month’s notice for preview deprecation. The launch announcement also says reinforcement-learning training is continuing. Consequently, keeping the same model name and prompts does not establish that two runs used an unchanged checkpoint.
The model-specific usage example calls mistral-large-4 with reasoning_effort set to high. Record the setting explicitly; the example does not establish the default. Do not substitute a generic “latest” alias and assume it selects this preview.
A proposed pilot should answer one workload question at a time:
Coding: use a fixed repository revision and acceptance tests controlled separately from generated changes. Count accepted fixes, regressions, retries and review effort. For migration work, the RohitAI parity-harness guide explains how to preserve reference behavior; it is not a Large 4 test.
Documents and images: check extracted fields and supporting passages against authorized reference material. Include missing or ambiguous information, so an apparently complete answer cannot hide unsupported details.
Long context: increase input length while keeping answerable questions and known evidence locations. Compare retrieval accuracy and cost with a smaller-context baseline; advertised capacity alone does not establish reliable recall.
Keep the endpoint, requested and returned model identity where available, run time, reasoning settings, prompt and input hashes, test-harness version and failures together. Rerun the acceptance set after changes. Measure time and cost per accepted task, including retries and human correction, rather than only output speed.
The launch post reports coding, workflow and vision evaluations, including results attributed to Artificial Analysis and a coding-quality study with Surge AI. These can help shortlist workloads, but their scores are not a production success rate for your application. No benchmark is independently reproduced in this guide.
Budget against the current Standard price display
At the October 6 check, Mistral’s Standard pricing table displayed $0.68 input, $0.07 cached input and $2.09 output per million tokens.
The model page crosses out $1.36, $0.14 and $4.18 beside those lower amounts. The launch card still displays the higher input/output figures. This is evidence of original-versus-current price presentation, not proof of a Standard-versus-Priority difference. The reduction’s duration and account-specific applicability were not established.
Illustrative token budget, using the displayed Standard rates:
No cache hits: 10 million input tokens plus 2 million output tokens cost 10 × $0.68 + 2 × $2.09 = $10.98.
Assumed 90% input cache-hit share: the same totals cost 1 × $0.68 + 9 × $0.07 + 2 × $2.09 = $5.49. This assumes 9 million input tokens are billed as cached; it is not an observed saving.
Both examples exclude tools, infrastructure, tax, regional surcharges and retries beyond the stated token totals. They use rounded displayed rates, so they are not invoice predictions. Check the chosen service mode and current account terms before committing a budget.
Mistral’s caching documentation says matching prefixes can be reused but cache hits are not guaranteed. It reports cached usage through usage.prompt_tokens_details.cached_tokens. Use actual usage records to replace the assumed hit share; do not budget every repeated prompt as a cache hit.
Private deployment needs more than model quality
The model page leaves approximate GPU RAM unspecified. Its “Open” label is not a substitute for the pending license. Active-parameter count alone is not a hardware-sizing guide.
Before committing to self-hosting, require a released repository and revision, license text, tokenizer and configuration, supported runtime and precision format, and capacity guidance for the intended context and concurrency. Test that actual package when available; a hosted preview result cannot certify a future local deployment.
Regional hosting is a separate decision. Mistral’s regional-inference documentation says model availability varies by region and excludes Agents, Batch and Files APIs from regional endpoints. It also specifies a 1.1× list-price multiplier. This guide has not verified Large 4 availability on a particular regional endpoint.
Therefore, evaluate the intended endpoint and required features together. A working agent on the general service does not prove the same application can move unchanged to a regional endpoint. Regional inference also does not locate all control-plane metadata in that region; processing location and retention need separate checks.
For the broader procurement decision, the RohitAI Cloudera–Mistral private-AI guide compares hosted and customer-operated routes. It does not establish Large 4 compatibility with Cloudera.
The immediate choice is narrow: evaluate hosted workload fit now, or wait when private operation is essential. Revisit procurement when the promised weights, actual rights and supported deployment requirements can be inspected.
Methodology: AI-assisted reporting and analysis based on linked public sources, checked on October 6, 2026. No authenticated API access, model benchmark or private deployment was tested. The evaluation plan and cost examples are proposed analysis, not observed results.
