Article

Fugu Max Turns Model Orchestration Into a Spending Policy

Sakana’s Fugu Max and Ultra v2 turn learned routing into an API product. Builders should measure accepted-task cost, drift, and opacity.

A Fugu-shaped coordinator routing requests across multiple AI models toward separate cost and quality targets

Sakana AI launched Fugu Max v1.0 and Fugu Ultra v2.0 on September 11, with both available immediately through its model API. The obvious reading is that Sakana added a cheaper Fugu and upgraded its strongest one. That is accurate, but it misses the more consequential change: Sakana is turning learned orchestration into a product with explicit operating points. Max is optimized for cost-performance; Ultra is optimized for answer quality. The coordinator is no longer just choosing which model should think. It is implicitly deciding how a customer’s inference budget should be spent.

RohitAI’s June analysis already argued that Fugu made orchestration feel like the model. The July Fugu-Cyber follow-up then showed why verification, hidden fan-out, and internal billing matter in a specialized workflow. September is not a rediscovery of multi-agent coordination. It is the first serious test of whether that coordination can become a general-purpose spending policy.

That makes Fugu more useful and harder to evaluate. Sakana’s launch charts plot quality against output-token list price. A buyer needs a different denominator: total cost per accepted outcome, including input, tools, retries, internal orchestration where billed, latency, and human review. Until that ledger exists, “frontier performance at a lower price” is a promising hypothesis, not a procurement conclusion.

The September thesis is not that a team of models can outperform one model. It is that a hidden coordinator can become a good steward of somebody else’s money. That claim must be tested at the task boundary, not read from a token-price chart.

What changed since June—and what did not

The product line now has two clearer economic targets. Fugu Max v1.0 is the new cost-performance endpoint. Fugu Ultra v2.0 replaces the previous quality-oriented Ultra alias. Sakana documents one-million-token context configurations, text and image input, and high or xhigh effort for both. The stable aliases currently resolve to fugu-max-v1.0 and fugu-ultra-v2.0, while explicit version IDs are also available in the model documentation.

Endpoint

Published operating point

Current list rates per 1M tokens

Cost caveat

Fugu Max v1.0

Cost-performance

$2 input / $6 output / $0.25 cached input

No context-length surcharge is listed. Internal web search and fetch calls cost $0.007 each. Max’s complete internal usage accounting is not yet clear.

Fugu Ultra v2.0

Maximum answer quality

$5 input / $30 output / $0.50 cached input; above 272K context, $10 / $45 / $1

Additional orchestration input, cached-input, and output tokens are billed at the corresponding rates.

These are confirmed rate-card facts from Sakana’s current pricing page, not estimates from the launch graphic. The subscription plan named “Max” is separate from the model named Fugu Max.

What did not change is equally important. Customers still call a model-shaped service while Sakana chooses the collaboration pattern behind it. Max and Ultra use fixed pools, and Sakana does not reveal the selected worker identities or coordination route for each request. Standard Fugu supports provider exclusions; Max and Ultra do not offer that control in the public model documentation.

The research lineage is real. The linked report is not new.

Fugu’s underlying idea did not appear overnight. Sakana’s Conductor paper describes a 7B coordinator trained with reinforcement learning to invent communication structures and focused instructions across randomized worker pools. It can select itself recursively as a worker. The TRINITY paper takes another route: it evolutionarily optimizes an approximately 0.6B language-model backbone and roughly 10,000-parameter routing head to assign Thinker, Worker, and Verifier roles over multiple turns.

Those papers establish that learned coordination is more than a hand-written router. They do not disclose Max’s or Ultra v2’s production parameter counts, worker budgets, training objective, or stopping policy. The technical report linked from the product pages was submitted in June and last revised on June 23. Its arXiv “v2” is a paper revision, not a September Ultra v2 report. Treat it as architectural history, not an evaluation appendix for the new endpoints.

task value and constraints
  -> choose Max or Ultra operating point
  -> hidden coordinator selects workers and roles
  -> tools, retries, and verification consume work
  -> final answer crosses the API boundary

buyer KPI = accepted outcomes / total task cost

RohitAI’s interpretation: the product’s durable asset may be the policy that qualifies workers, assigns roles, and decides when more inference is worth buying—not exclusive ownership of every worker model. That is strategically attractive because worker models can change. It also creates a new switching cost: customers become dependent on a coordination policy they cannot inspect or reproduce.

A cheaper output token is not a cheaper completed task

Fugu Max’s output rate is 40% below Sonnet 5’s $10, 50% below GPT-5.6 Terra’s $12, and 60% below Kimi K3’s $15. Those statements are arithmetically true. They compare one tariff line. They do not show what an application pays for the same accepted result.

Consider an intentionally simple illustration: every system is billed for exactly 100,000 uncached input tokens and 10,000 output tokens. This is not the same text, an observed run, or a forecast of savings. It merely combines current unit prices.

Model

Input / output per 1M

Illustrative bill

Max difference

Fugu Max v1.0

$2 / $6

$0.26

Baseline

Sonnet 5

$2 / $10

$0.30

Max is 13.3% lower

GPT-5.6 Terra

$2 / $12

$0.32

Max is 18.75% lower

Kimi K3

$3 / $15

$0.45

Max is 42.2% lower

The input share compresses the apparent advantage. Real tasks add more variables: cached context, tokenizer differences, tool calls, retries, review, and possibly different output volumes. Ten Max web search or fetch operations add $0.07 before tokens—more than the $0.04 gap in the equal-volume Sonnet illustration. Ultra is clearer but more complex: its internal orchestration categories are additional billable work, and its rates step up above 272,000 context tokens. Its final-output limit constrains the answer, not the coordinator’s internal work.

The uncomfortable gap is Max accounting. Sakana explains Max’s public token and tool rates, but the documentation reviewed for this release does not establish whether every form of hidden coordination is surfaced or billed in the same way as Ultra. Do not assume it is free. Ask for an invoice-level field definition, then reconcile sample usage objects against the bill before passing costs through to customers.

The benchmark charts are a trial invitation, not a verdict

Sakana reports that Max leads the selected comparison on six benchmarks and extends the plotted cost-performance frontier on seven of ten. It reports Ultra v2 as best or joint-best on five of eight tests and top-two on seven. These are first-party launch results. Sakana’s charts do not publish enough matched information about versions, effort, tools, retries, judges, trial counts, or confidence intervals to establish independent superiority.

Signal

Company-reported result

Evidence-aware read

Max on Terminal Bench 2.1

89.5, 0.1 above the plotted Gemini 3.8 Flash result

A tiny chart margin without matched uncertainty should be treated as competitive, not decisive.

Ultra v2 on Chartography

48.3

Surge’s visible board lists Fable 5.1 Max at 46.2 and GPT-5.6 Sol Max at 45.0, but no Fugu row. The current relevant gap is 2.1, not the launch’s 18.8-point comparison with older Fable 5 High.

Ultra v2 on DeepSWE

74.3

Datacurve’s September 3 board has leading systems around 74% with uncertainty intervals and no visible Fugu entry. Similar magnitude is not replication.

Ultra v2 on Toolathon

80.6 versus 75.0 for Ultra v1.1

The 5.6-point uplift is a useful reason to test multi-tool workflows, not proof of the mechanism that produced it.

Ultra v2 on GPQA-D

95.5 versus 95.6 for Ultra v1.1

The upgrade is uneven. It should not be marketed internally as a universal quality increase.

The external benchmark context strengthens the case for pilots while narrowing the claims. Chartography has 100 professional-chart tasks and evaluated 30 configurations with 20 scored trials per task. DeepSWE contains 113 original long-horizon engineering tasks across 91 repositories and five languages. Those are meaningful workload probes. They are not broad measures of intelligence, and the benchmark owners’ public boards do not currently reproduce the new Fugu entries.

RohitAI’s interpretation: the uneven change from Ultra v1.1 is more useful than the leaderboard rhetoric. A 5.6-point Toolathon gain, a 1.3-point Chartography gain, and a 0.1-point GPQA-D decline suggest that v2 may be a workflow-specific upgrade. Builders should separate tool orchestration, professional visual reasoning, long-horizon coding, and short hard questions instead of averaging them into one “smartness” number.

A diverse pool can still create one opaque dependency

Sakana’s Max pool illustration includes closed and open models, NVIDIA Nemotron, Sakana Namazu, and recursive Fugu Max. The company does not disclose a complete inventory, exact pool size, or which workers handled a given query. “Includes open models” therefore does not mean “all-open service,” and worker diversity does not make the Sakana endpoint itself redundant.

This leads to a non-obvious business result. Open-weight workers can strengthen a closed coordination product. If capable workers become interchangeable inputs, value moves toward worker qualification, routing, communication, and stopping. Sakana can reduce dependence on any one worker supplier while increasing the customer’s dependence on Sakana’s private policy. Supplier concentration goes down inside the pool; orchestrator concentration goes up at the API boundary.

The missing resilience metric is replacement regret: how much accepted-task quality, latency, and cost deteriorate when an important worker is removed or materially changed. Sakana’s Gemma conductor experiment is useful research evidence that the coordinator layer can be retrained across backbones. It is not a production worker-removal test for Max or Ultra v2. Because customers cannot alter those fixed pools, a meaningful removal study requires Sakana’s cooperation or contractual evidence.

A long worker list is inventory. Resilience is the measured degradation after one of the important workers becomes unavailable.

Alias drift is now part of model risk

The alias fugu-ultra now resolves to Ultra v2.0. That sounds routine until effort semantics are included. Sakana’s getting-started catalog documents Max and Ultra v2 with high and xhigh effort, with max acting as an alias for xhigh. Ultra v1.1 had a distinct maximum effort level. An application can therefore keep the same model alias and effort string while the effective behavior changes beneath it.

RohitAI’s interpretation: semantic alias drift is a more dangerous reproducibility trap than a visible benchmark update. Pinning the version ID makes the requested product explicit, but Sakana has not published a guarantee that a version ID freezes every worker snapshot and routing policy. Record the returned model, effort, date, usage object, application revision, and eval-set revision. For regulated or high-value work, ask what the version contract actually freezes.

Protocol compatibility does not solve this. Fugu supports Responses, Chat Completions, Models, and an Anthropic-compatible Messages endpoint, but previous_response_id is unsupported. Clients must resend history. The service accepts but ignores temperature and parallel_tool_calls. The server controls tool parallelism. A familiar request shape lowers integration work; it does not promise identical state, sampling, latency, or recovery behavior.

The outer agent runtime still belongs to you

Fugu is a coordinator exposed as a model. It is not a durable application runtime. The inner system can decide which workers should reason; the outer application must still own conversation history, tool permissions, checkpoints, retries, cancellation, acceptance tests, and the record of external side effects. This distinction matters in coding tools, where Sakana’s public repository offers Codex and Claude Code integration but the surrounding harness still determines what the agent can inspect, change, and approve.

The June technical report described isolated worker contexts inside a workflow with persistent tool history across workflows. That is an interesting design for reducing correlated reasoning, but it is historical evidence, not proof that every v2 mechanism is unchanged. A useful customer test is to compare independent reviewers that do not see the first answer with reviewers that do. If every verifier inherits the same mistaken premise, agent count is not independent evidence.

The operational lesson matches RohitAI’s analysis of the OpenAI Agents API: managed intelligence can simplify execution without replacing the customer’s outcome ledger. Otherwise a harness improvement gets credited to the model, a model regression gets blamed on the harness, and neither vendor nor buyer can explain the bill.

A pilot that can answer the buying question

RohitAI did not run paid Max or Ultra v2 evaluations for this article. The responsible next step is a controlled pilot, not a production recommendation. Use real work that is hard enough to justify orchestration and structured enough to score.

  1. Pre-register workload cohorts. Separate repository changes, professional charts and PDFs, multi-tool workflows, and short factual tasks. Do not average them before inspecting each category.

  2. Pin fugu-max-v1.0 and fugu-ultra-v2.0, set effort explicitly, and freeze the competing model IDs, prompts, tool versions, price snapshot, and acceptance rubric.

  3. Hold the outer harness constant. Give every candidate the same repository state, tools, network rules, timeout policy, retry budget, and opportunity to recover.

  4. Capture the complete usage object and invoice. Record visible tokens, cached tokens, orchestration categories, tool-call counts, retries, p50 and p95 wall time, and cancellation behavior. Clarify Max field semantics before comparing bills.

  5. Judge outcomes, not prose. Use tests, numerical-grounding checks, source verification, or blinded reviewers. Count reviewer minutes and escaped errors; a persuasive wrong answer is not an accepted task.

  6. Set three independent limits: wall-clock time, external tool calls, and total spend. Ultra’s final-response token limit is not a full-work budget.

  7. Replay a stable slice after alias or pool updates. Ask Sakana for a provider-removal or substitution test if continuity is part of the purchase case. Measure drift instead of assuming diversity equals resilience.

Metric

Calculation

What it prevents

Accepted-task cost

Tokens + tools + retries + review cost, divided by accepted tasks

Mistaking a low output tariff for a cheap workflow

Escaped-error rate

Errors found after acceptance, divided by accepted tasks

Rewarding confident but brittle outputs

Reviewer burden

Median and p95 human minutes per accepted task

Hiding cleanup behind API spend

Replacement regret

Change in quality, cost, and latency after a required worker change

Treating pool size as continuity evidence

Behavioral drift

Outcome change on a frozen replay set after service updates

Letting stable aliases masquerade as stable systems

Retain a Fugu configuration only if it improves accepted outcomes per dollar without an unacceptable increase in p95 latency, reviewer burden, or supplier constraints. A 0.1-point vendor-chart lead is not a rollout threshold.

Commercial constraints belong in the eval

Sakana’s public terms of service are part of the product boundary. They exclude access from the EEA, United Kingdom, and Switzerland. They permit training use and human review of content, with a training opt-out that is not retroactive. They also restrict using the service or its output to develop another model or a competing AI service. Enterprise arrangements may differ, but this research did not establish any exceptions.

Before moving proprietary code, customer documents, or regulated data, resolve the contract rather than infer protections from a toggle. Training use, retention, human review, subprocessors, worker-provider access, residency, deletion, and breach response are separate questions. A fixed multi-provider pool is especially relevant when a customer must know every recipient of its prompts.

  • Confirm that the account region and intended use are eligible before building around the endpoint.

  • Obtain written answers for training, retention, human review, subprocessors, residency, and deletion before sending sensitive material.

  • Distinguish standard Fugu provider exclusions from the fixed Max and Ultra pools.

  • Review the competing-service restriction if your product trains, evaluates, routes, or serves other models.

  • Keep application-owned history and checkpoints because the API does not support server-side continuation through previous_response_id.

Where Max, Ultra, and a direct model each fit

Try Fugu Max first

Use Max as a candidate for evaluable, moderately high-value work where output price matters but a single cheap model misses too often: repository maintenance, grounded research with bounded retrieval, professional document extraction, and tool workflows with clear acceptance tests. Its strongest buying case is not “cheapest model.” It is “fewer failed attempts at a tolerable total bill.”

Escalate selectively to Ultra v2

Use Ultra for difficult tasks whose expected value can absorb additional internal work: complex code changes, high-stakes analytical documents, multi-stage tool use, or cases where one better answer can remove substantial review. Keep spend, tool, and wall-clock controls outside the model. Long context above 272K changes the rate card, and a short final answer does not prove the run was cheap.

Keep a direct model path

A direct model remains the better control—and may remain the better product—for low-latency chat, simple extraction, stable classification, strict provider provenance, contractually restricted data, or workflows where internal fan-out adds little value. Maintain that path even if Fugu wins the pilot. It preserves a baseline, a fallback, and leverage in future procurement.

Three predictions worth tracking

  1. Orchestration vendors will expose more operating points. Max and Ultra split cost-performance from quality, but real workloads also optimize for latency, provenance, privacy, and bounded spend. Expect tiers or controls built around those constraints over the next 6–12 months.

  2. Enterprise buyers will ask for policy manifests rather than full traces. Sakana is unlikely to reveal its proprietary route for every request, but customers can reasonably demand versioned worker classes, supplier-exclusion guarantees, change notices, and replay evidence.

  3. Credible orchestration evaluations will add accepted-task cost and worker-removal tests. Score-versus-output-price charts are easy to publish. The harder and more decision-useful evidence is whether quality and economics survive retries, reviewers, and a changed worker pool.

These are RohitAI forecasts, not Sakana commitments. The common thread is that the coordinator is becoming a governable economic layer. The market will eventually demand evidence at that layer, not only benchmark scores from the answers it emits.


Frequently asked questions

Is Fugu Max a newly trained frontier foundation model?

Not in the usual monolithic sense. Sakana presents Max as a learned orchestrator that coordinates a pool of worker models behind one endpoint. The company has not disclosed a September Max checkpoint, parameter count, or complete worker manifest.

Is Fugu Max cheaper than Sonnet 5, GPT-5.6 Terra, or Kimi K3?

Its $6-per-million output rate is lower than those current output rates. That does not establish lower task cost. Input volume, output volume, cache, tools, retries, hidden coordination, latency, and review can change the result. Compare the same accepted tasks and the complete bill.

Does Sakana reveal which models answered a request?

No. Sakana says selected worker identities and coordination routes are proprietary. The public materials describe a mixed pool and identify some families, but not a per-request route or complete manifest.

Are the Max and Ultra v2 benchmark results independently replicated?

Not in the public evidence reviewed for this article. The launch scores are first-party. Benchmark owners provide useful task definitions and independent boards, but visible Chartography and DeepSWE boards did not contain a Fugu result that reproduced Sakana’s setup.

Does OpenAI compatibility make Fugu a drop-in replacement?

It reduces protocol work, but behavior differs. Server-side previous-response continuation is unsupported, some accepted controls are ignored, tool parallelism is server-controlled, and an orchestrated response can have different latency and accounting. Run integration and recovery tests before migration.

Should builders use the aliases or explicit version IDs?

Use explicit IDs during evaluation and for reproducibility-sensitive production paths. Aliases are convenient for automatic upgrades, but the Ultra transition shows that model and effort semantics can change while application code stays the same. Keep a replay set for either choice.

Does the one-million-token context window have one price?

No. Sakana currently lists no context-length surcharge for Max. Ultra v2 steps from $5/$30/$0.50 to $10/$45/$1 above 272K context, and its additional orchestration tokens are billed at the corresponding rates. The one-million-token catalog value is a documented configuration, not evidence that every task uses that context reliably.

The useful way to read this launch

Fugu Max and Ultra v2 make Sakana’s bet much more concrete. June established the technical proposition: a learned coordinator can assemble model workers behind one endpoint. September adds a commercial proposition: that coordinator can offer distinct quality-and-cost operating points and keep improving as the worker market changes.

That proposition is credible enough to test and too opaque to accept on launch charts alone. Max’s output rate is genuinely attractive. Ultra v2 shows meaningful company-reported gains on some tool and professional-work tasks. The API is familiar. But the buyer still lacks a full Max accounting model, independent replication, a worker manifest, a published pool-freeze contract, and evidence of low replacement regret.

The right response is neither dismissal nor automatic migration. Give the orchestrator work whose outcome you can score. Pin the version. Preserve the outer runtime and direct-model baseline. Count every token, tool call, retry, minute, and rejected result. If Fugu wins on accepted outcomes per dollar under those conditions, Sakana has built something more valuable than another model endpoint: a spending policy worth renting.