Article

Nemotron IOI and IMO: Models, Datasets and How to Start

Choose NVIDIA’s Nemotron IOI and IMO artifacts, check compute and context requirements, and plan a bounded proof-workflow run without overstating the scores.

Editorial illustration for Nemotron IOI and IMO: Models, Datasets and How to Start: a controlled task flows from input to output. Not documentary evidence.

On October 7, 2026, NVIDIA published a roundup of its Nemotron IOI and IMO research, connecting competitive-programming and olympiad-proof systems with available weights, datasets and inference recipes. For coding-agent builders, reasoning researchers and evaluation teams, the practical choice is which component to reuse—not whether a medal-level score makes a model ready for their product.

The roundup is new; the underlying work is older. The coding paper first appeared on September 2 and the math paper on September 9. The model cards date their releases to September. Treat this as an implementation guide to the available artifacts, not an October 7 launch of every checkpoint.

Choose the artifact before choosing the cluster

Start with evaluation data if your question is whether a model can solve difficult proofs. Start with training traces if you are studying better generators or critics. Provision the specialist models only when you need to investigate their behavior or the published inference system.

Artifact

What is available

Useful starting point

Nemotron-IMO-Bench

200 problems with reference proofs; 50 each in algebra, combinatorics, geometry and number theory.

Compare proof systems. Keep reference proofs out of generator inputs and exclude the benchmark from training.

Nemotron-Math-Proofs-v3-SFT

414,890 records across 15,818 problems; 121.95 GiB in ten JSONL shards.

Study proof generation, refinement, verification and meta-verification. Preserve task labels and split by problem identity.

Nemotron-Math-Proofs-v3-RL

9,597 proof-generation prompts, formatted for NeMo-Gym. No completed policy outputs or realized rewards.

Supply a policy, verifier and trainer; the prompt file is not a completed RL run.

Ultra Math SFT and RL checkpoints

Two specialist checkpoints, alongside the general Ultra model used in the ensemble.

Compare specialists or configure the three-model proof workflow; these are not small-model downloads.

Competitive-Coding 550B-A55B-NVFP4

A competitive-programming specialist with 550B total and 55B active parameters; up to 262,144 tokens of context.

Study algorithmic code generation and feedback-driven refinement, not general coding-assistant readiness.

The cards label the proof datasets and benchmark CC BY 4.0, and the specialist weights OpenMDW-1.1. Check each artifact’s terms and revision separately. The IOI paper says the full coding training corpus cannot be distributed because of third-party restrictions; available weights and a described training procedure do not remove that reproduction limit.

For the SFT corpus, metadata.data_type distinguishes the four tasks. Group related traces by problem before making training and validation splits, so a critique of a held-out proof does not leak into training. For the prompt-only RL release, the guide to tasksets, runtimes and reward audits explains the separate responsibilities.

Read the competition results as system results

  • IOI: NVIDIA reports 535.4/600 from an unofficial, unsupervised run outside the official rankings. Its paper describes matching contest time, internet-access and submission constraints, with a peak allocation of up to 760 GB300 GPUs. That is evidence about a specialist plus a substantial search system—not one ordinary model response.

  • IMO: The authors report 30/42, a gold-level score, and say official IMO graders assessed the submitted proofs. This is the research team’s account of the result; no independent reproduction was established.

The IMO recipe implements a generate-verify-refine workflow: propose proofs, critique them, revise candidates and select a final submission. It uses natural-language judgments, not a formal proof checker. In the IMO report, model graders estimated roughly 32 points at the deadline, versus the reported official 30. Agreement among model critics therefore needs separate correctness review.

Budget serving, context and verification separately

The coding card’s serving example uses four GB300 GPUs. The math RL card recommends eight B200 GPUs for a single-node BF16 deployment. Neither is a universal minimum, and neither describes the resources needed to serve all three proof checkpoints concurrently at full search throughput.

There is also a concrete configuration mismatch to resolve: the math card starts its server at 262,144 tokens, while the pinned full IMO configuration expects 524,288-token contexts and completion budgets as high as 512,000. Prompt tokens consume context too. Reconcile server limits, tokenizer-based budgeting and concurrency before launching; shortening the limits changes the experiment.

The full configuration allows 384 initial proof attempts and 16 verification judgments per distinct proof. Calculated upper bound: 384 × 16 = 6,144 verification requests per problem in round one. This assumes every attempt yields a distinct eligible proof and no early cancellation. It excludes later refinement, final judging and retries; it is not a measured call count or a token bill.

The smoke configuration cuts the first round to six attempts, verification to two judgments per proof and final judging to three per finalist. It uses two rounds and 131,072-token contexts, but retains three checkpoint roles. Fewer verifier votes change the acceptance procedure as well as cost; a smoke run cannot inherit the full system’s score.

The checked materials do not establish an all-in reproduction price or a publicly priced hosted service for all three exact math checkpoints. Confirm provider access separately, then price the chosen deployment after fixing checkpoint, context, concurrency and search budgets.

A bounded first run of the IMO recipe

The following is a documentation-derived setup path, not an executed tutorial. Use a NeMo-Skills checkout with its required dependencies installed. The inspected recipe revision is bcf059af55c20a89f797724598f9908d126153e6; pin model and dataset revisions separately.

  1. Prepare one or two problems. The recipe accepts JSONL with a unique problem_idx or id, and a statement in problem or question. Keep reference answers on the evaluation side. If using Nemotron-IMO-Bench, distinguish any problems used in the report’s development subset from held-out evaluation.

  2. Configure the smoke file. Copy recipes/nemotron-imo-tts/configs/smoke.yaml to my-run.yaml. Fill the input path, endpoint and served model names for the general, RL and SFT roles. Check actual model access and effective context limits before inference. Set concurrency to the capacity of your deployment.

  3. Validate before spending on inference. With configuration complete, run the documented dry-run command below from the checkout. It writes manifests and validates configuration without issuing inference requests; tokenizer loading and endpoint model discovery may still occur.

python recipes/nemotron-imo-tts/run.py --config my-run.yaml --output-dir runs/my-run --dry-run

Before removing --dry-run, set a bounded compute budget and review the configuration. After an actual run, inspect submissions.jsonl, results.jsonl and errors.jsonl. An empty submission or no_candidates is not a solved problem. Use a fresh output directory when changing experiment settings; preserve manifests and request records.

Treat integration success and proof quality as separate outcomes. Independently review the selected proof, record truncations and failures, and account for generation, verification and refinement. Only then enlarge the problem set or search budget. Replacing the three specialists with an accessible model can explore the pattern, but it is an adaptation rather than reproduction of NVIDIA’s result.

For coding teams, evaluate the program users receive

The coding card provides a serving example and the IOI paper explains GenCorrect. This review did not establish a single pinned end-to-end command reproducing the exact 2026 live run, so those materials should not be presented as a one-command replication.

For an adaptation, use representative executable tasks and an isolated runner. Track final-program correctness, candidate and submission counts, wall time, compute consumption and failures. Keep feedback used during refinement separate from the final held-out tests. If your product reviews pull requests, the ReviewBench evaluation guide addresses that reader task more directly than a competition score.

Start by inspecting the benchmark, data schema or configuration that answers your research question. Commit to a larger run only when its outputs and total cost will inform a concrete model or workflow decision.

Methodology: AI-assisted reporting and analysis using NVIDIA’s papers, model and dataset cards, and repository files rechecked on October 7, 2026. The request-count estimate is arithmetic on the published configuration. No model inference, training, deployment test or competition reproduction was performed.