One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO
One Model Family, Two Gold Level Results: Fine Tuning Nemotron for IOI and IMO The International Olympiad in Informatics (IOI) and the International Mathematical Olympiad (IMO)...
By AI Engineering Team
One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO
The International Olympiad in Informatics (IOI) and the International Mathematical Olympiad (IMO) evaluate different abilities. IOI requires contestants to design algorithms and submit code that passes hidden tests within strict time and submission limits. IMO requires rigorous proofs written in natural language. Achieving gold-medal-level performance in both competitions requires broad specialization.
Results from 2026 show how Nemotron can serve as a foundation for specialist models. Starting with Nemotron 3, NVIDIA teams used supervised fine-tuning (SFT), reinforcement learning (RL), and feedback-driven inference to develop systems that reached gold-medal level at both IMO 2026 and IOI 2026.
| Competition | Nemotron specialization | Result |
|---|---|---|
| IOI 2026 | Nemotron-3-Ultra-CC with SFT and GenCorrect | 535.4/600, above the 361.12 gold threshold and the top human score of 498.27 |
| IMO 2026 | Nemotron 3 Ultra general, SFT, and RL checkpoints in a generate-verify-refine system | 30/42, above the official gold threshold of 29 |
The IOI result came from a live, prospective run under the same time, internet-access, and submission constraints as human contestants. It was an unofficial, unsupervised benchmark and was not included in the official IOI ranking. Official IMO graders evaluated the proofs submitted by the IMO system.
A reusable specialization recipe
A model being easy to fine-tune means more than simply providing a trainable checkpoint. A capable foundation model should be adaptable to demanding domains through a clear and reusable process.
The two projects followed four main steps:
- Start with a strong Nemotron base model.
- Curate domain-specific problems and high-quality reasoning traces.
- Apply post-training methods such as SFT and, where useful, RL.
- Pair the specialist model with an inference loop that generates, evaluates, and improves candidate answers.
The training and inference runs were substantial, but the overall approach used established methods. A new foundation model was not required for each challenge. Instead, Nemotron was specialized for each task.
From general coding ability to IOI gold
For competitive programming, the team curated 22,000 problems and generated synthetic reasoning traces to train two specialist models. Nemotron-3-Nano-CC has 30 billion total parameters and 3 billion active parameters, and received both SFT and RL. Nemotron-3-Ultra-CC has 550 billion total parameters and 55 billion active parameters, and received SFT.
Results from IOI 2025 illustrate the effect of specialization. Nano improved from 130 points before post-training to 280 points after SFT and 291 points after RL. With GenCorrect, an iterative generate-evaluate-refine strategy, it reached 468 points, exceeding the gold threshold of 438.3. Ultra-CC reached 502 points with the same test-time strategy.
These experiments also showed that adaptation can vary by model scale. SFT produced most of Nano's improvement, while RL added a smaller but consistent gain. For the stronger Ultra model, one SFT epoch was enough to outperform the fully post-trained Nano model across IOI, ICPC, and LiveCodeBench Pro. This result informed the competition-specific Ultra-CC system used for IOI 2026, which scored 535.4 out of 600.
Teaching Nemotron to prove, check, and revise
The IMO project applied the same general approach to olympiad mathematics. Starting from Nemotron 3 Ultra, the team trained one specialist with SFT and another with RL.
The SFT corpus contained 414,890 quality-filtered examples covering 15,818 unique proof problems. The data focused not only on final answers, but also on proof generation, refinement, verification, and meta-verification. This enabled the model to construct arguments, identify gaps, respond to critiques, and assess whether a proof was complete. The RL model was trained on 9,597 proof problems selected near the model's capability frontier.
Both post-trained checkpoints outperformed the general-availability model during development experiments. The SFT checkpoint was strongest in the first search round, while the RL checkpoint achieved the best overall single-checkpoint result. Because their strengths were complementary, the final system used both specialists together with the general model.
For each IMO problem, the models generated candidate proofs, scored them, created critiques, and refined the most promising attempts. A separate high-compute stage selected the final submission. The system operated entirely in natural language, without a formal prover, external tools, or internet access. It scored 30 out of 42 points, including full credit on four of the six problems, and exceeded the official gold-medal threshold.
Fine-tuning and test-time compute work together
Earlier work on IOI 2025 showed how test-time compute can improve the performance of open-weight models. The 2026 results add that specialization can provide an inference system with better candidates, critics, and refinements.
At IOI, GenCorrect converted the gains from fine-tuning into larger improvements across multiple feedback rounds. At IMO, combining complementary SFT and RL checkpoints was more effective than simply drawing more samples from one checkpoint. In both projects, the strongest results came from combining a capable specialist with an inference system that could search, verify, and improve answers.
The results therefore depended on more than fine-tuning or brute-force sampling alone. They came from jointly designing the model, the data, and the inference loop.
Models, data, and recipes
The Nemotron Labs IMO 2026 collection includes the SFT and RL checkpoints, both training datasets, and Nemotron-IMO-Bench, a benchmark containing 200 olympiad-level problems. The IMO paper describes the training approach and generate-verify-refine system. The NeMo-Skills repository includes the IMO inference pipeline, prompts, submitted proofs, and a reproducible quickstart.
For competitive programming, the Nemotron-3-Ultra-CC model is available on Hugging Face. The IOI paper describes its training recipe and the GenCorrect methodology, while the IOI evaluation and inference pipeline are also available in NeMo-Skills.
Together, the IMO and IOI results demonstrate that Nemotron can be fine-tuned into specialized models for demanding domains and combined with inference workflows that generate, verify, and refine solutions to problems at the level of human competition.