On September 9, 2026, Mistral reported migrating 40,000 lines of Fortran 77 to C++ for an unnamed European energy operator’s reservoir simulator. For teams maintaining numerical legacy systems, the useful question is how to establish that an agent-generated replacement preserves the behavior they need.
A parity harness compares the old and new implementations on declared inputs and checkpoints. This guide proposes how to specify that comparison, protect its reference data and decide whether a bounded migration pilot is ready to begin.
What the case establishes—and leaves unknown
Mistral describes a runnable baseline without a test suite or centralized documentation. It exported final results and engineer-selected intermediate states into a C++ comparison framework. After autonomous translation retained legacy structure and a structured agent attempt stalled, the team used human-supervised coding, testing and review. Domain engineers reviewed the proposed architecture; humans accepted the pull requests. These are Mistral’s reported project observations, not an independent benchmark.
The first sprint covered 40,000 of 300,000 source lines. The account does not establish full-project completion, total duration, cost or a controlled productivity gain. Its week-long attempt was not the complete delivery timeline. Those gaps prevent a reliable outsourcing or delivery-time comparison.
Separate three acceptance decisions
The following is a proposed review framework. A numerical match answers only one part of the modernization decision; give each part an owner and a recorded outcome.
Decision | Evidence to require | What a pass does not establish |
|---|---|---|
Behavior preserved | Required cases and checkpoints pass the approved comparison rules. | Correctness on untested inputs or failure paths. |
Design improved | A maintainer reviews interfaces, state ownership, dependencies and integration. | Numerical fidelity merely because the new code looks cleaner. |
Domain requirements met | A domain owner checks assumptions, invariants and intended changes. | Physical or business validity merely because outputs match the old system. |
Keeping migration fidelity separate from robustness also has research support. A GAMESS modernization preprint studies paired fault injection because ordinary successful runs do not cover behavior under perturbed internal state. That is a specialist extension for relevant systems, not a requirement to reproduce that experiment for every pilot.
Build a harness that can reject the wrong migration
1. Preserve an explainable baseline
Create a manifest identifying the source revision, binary, compiler and options, numerical libraries, runtime environment and input files. Keep their hashes with the captured results. Run representative cases repeatedly to characterize existing variability before choosing comparison bounds; a changing baseline makes a discrepancy harder to diagnose.
Name the maintainer and domain owner who can explain the outputs. Record known defects and excluded cases. If the reference cannot run, fund baseline recovery first. If it runs but nobody can judge whether its behavior is acceptable, limit the pilot to documentation and investigation.
2. Define every checkpoint before comparing numbers
For each observable, specify its meaning, units, capture point, dimensions, element order, precision and handling of missing or non-finite values. Declare whether it needs exact equality or an approved numerical tolerance. Keep protected reference checkpoints separate from candidate outputs, and send both through the same comparator before human acceptance.
Representation needs deliberate mapping. GNU Fortran’s interoperability documentation explains that Fortran A(n,m) corresponds to C A[m][n], with different default index origins. A checkpoint loader that compares the wrong elements can invalidate an otherwise sensible tolerance policy.
Also verify that adding checkpoint instrumentation leaves the ordinary baseline results unchanged. A checkpoint collected from an altered reference is not automatically trustworthy.
3. Make tolerance a domain decision
For finite values, one possible element-wise policy is the NumPy isclose rule:
abs(actual - reference) <= atol + rtol * abs(reference)Here, atol is an absolute allowance in the quantity’s units; rtol scales with the reference magnitude. The comparison is asymmetric. NumPy warns that its default absolute tolerance can be unsuitable near zero. Have the domain owner approve bounds per quantity; do not increase them simply because a candidate fails.
Library defaults need an explicit decision too. NumPy testing.assert_allclose defaults to equal_nan=True and strict=False. Matching NaNs can therefore pass, and scalar-versus-array comparisons receive special treatment. Setting strict=True requires matching shape and dtype and disables that scalar behavior. Add an explicit finite-value check wherever NaN or infinity is forbidden.
NumPy is an illustrative comparator option, not a requirement or a description of Mistral’s C++ harness. If a dtype change is intentional, specify the conversion and its acceptance policy separately rather than silently discarding a type mismatch.
4. Prove that failures remain visible
Before using the harness to accept code, introduce negative controls: a value outside the approved tolerance, a missing checkpoint, the wrong shape and an unexpected NaN. Each must be rejected for the intended reason. Record executed, skipped and failed cases separately; a missing required test is not a pass.
Protect acceptance fixtures and thresholds from routine migration edits. An agent can propose a reference change, but require separate review of why the expected behavior changed. A second reviewer agent does not provide independent evidence if both agents can rewrite the same expected answers.
For related scientific workflows, our BootLoops verification guide examines the distinction between an executable check, its reference evidence and scientific judgment.
Run one bounded agent pilot
Choose a module with a reviewable input/output contract. Document shared state, external I/O, dependencies and side effects—not just its callers. Specify the desired interface before implementation. Keep intentional bug fixes separate from behavior-preserving translation, with their own expected results.
For Vibe, select the execution mode deliberately. The current permissions documentation says plan is read-only, while accept-edits approves edits but still asks about shell commands. Programmatic --prompt runs fall back to auto-approve if --agent is omitted; the interactive default_agent setting does not carry over. Pass --agent plan explicitly for read-only scripted exploration.
For implementation, use an isolated runner containing only approved code and test data, and inspect repository-provided agent configuration before trusting it. Record the CLI version, selected model and effective configuration. Mistral describes admin-managed configuration as settings distribution, not an enforceable security boundary; use infrastructure controls for isolation.
Bound runs with --max-turns and the required tool set, and retain machine-readable output where useful. The CLI documentation warns that --max-price uses potentially missing or outdated configuration prices: it is not a hard spending limit. Use provider-side budget controls where available.
Require a review packet: source and target identities, input and checkpoint hashes, comparator version, per-case results, mismatches and untested behavior.
Stop on an unexplained divergence, a required skipped case or repeated retries without a new diagnosis. Escalate to the maintainer or domain owner instead of relaxing acceptance.
Approve behavior, architecture and domain requirements separately. Add integration and error-path checks appropriate to the module, then define rollback and staged release before production use.
Price the accepted scope, not the generated code
As checked on October 5, 2026, Vibe’s product page offers Free, Pro, Team and Enterprise paths. The USD pricing page displays Pro at $14.99/month and Team at $24.99/user/month, excluding taxes and subject to fair-use limits. Free has limited coding sessions; Enterprise is contact-sales with private deployment options. These are access prices, not migration quotes.
Ask a vendor or services partner to price the same bounded module against the same acceptance rules. Keep a pilot ledger covering baseline recovery, harness preparation, documentation, model usage, solver compute, human review, rework and ongoing operations. Allocate staff time once so review and retry costs are not double-counted.
Require a named long-term owner and a handover package that can rerun the acceptance checks. Confirm deployment-specific data handling, entitlements and usage limits before supplying proprietary code. Neither a low subscription price nor a large generated diff establishes a lower total delivery cost.
The next decision is concrete: recover the baseline, resolve the acceptance specification, or start one bounded migration. Expand only after the pilot supplies reproducible accepted changes and usable cost records.
Methodology: This AI-assisted guide uses published Mistral material, official technical documentation and a research preprint, checked on October 5, 2026. The workflow is proposed engineering analysis. No migration, comparator test or independent reproduction of Mistral’s project was performed.
