Az NVIDIA bemutatta, hogyan alakította át Nemotron 3 alapmodelljeit olimpiai szintű feladatmegoldókká a 2026-os Nemzetközi Informatikai Diákolimpia (IOI) és a Nemzetközi Matematikai Diákolimpia (IMO) feladatai alapján. A fejlesztők felügyelt finomhangolást, megerősítéses tanulást és visszacsatolás-alapú következtetést alkalmaztak a siker érdekében.
Az informatikai feladatokhoz a GenCorrect nevű, iteratív generáló-értékelő-javító stratégiát használták, amellyel a Nemotron-3-Ultra-CC modell 535,4 pontot ért el a 600-ból. A matematikai bizonyításokhoz egy 414 890 példát tartalmazó adatbázison tanították be a modellt, amely külső eszközök nélkül, tisztán természetes nyelven ért el 30 pontot a 42-ből.
Az NVIDIA a teljes projektet nyílttá tette a közösség számára. A Hugging Face felületén elérhetővé tették a Nemotron Labs IMO 2026 gyűjteményt, a betanítási adatbázisokat, a Nemotron-3-Ultra-CC modellt, valamint a reprodukáláshoz szükséges kódokat és pipeline-okat is.
Az eredeti szöveg (Hugging Face)
A reusable specialization recipe From general coding ability to IOI gold Teaching Nemotron to prove, check, and revise Fine-tuning and test-time compute work together Open models, data, and recipes on Hugging Face The International Olympiad in Informatics (IOI) and the International Mathematical Olympiad (IMO) test different skills. IOI requires algorithms and code that pass hidden tests under strict time and submission limits. IMO demands rigorous natural-language proofs. Success at either competition is difficult. Success at both points to something broader.
Our recent results show that Nemotron is a strong, adaptable foundation for building world-class specialist models. Starting from Nemotron 3, our teams used supervised fine-tuning (SFT), reinforcement learning (RL), and feedback-driven inference to create systems that reached gold-medal level at both IMO 2026 and IOI 2026.
The IOI result came from a live, prospective run under the same time, internet-access, and submission constraints as human contestants. It was an unofficial, unsupervised benchmark and was not included in the official IOI ranking. The IMO system’s submitted proofs were graded by official IMO graders.
"Easy to fine-tune" should mean more than making a checkpoint trainable. It should mean that a capable foundation model can be adapted to a demanding domain with a clear, reusable recipe.
Across the two projects, that recipe had four parts:
The training and inference runs were substantial, but the underlying approach is familiar and reproducible. We did not need to build a new foundation model for every challenge. We specialized Nemotron for the task.
For competitive programming, we curated 22,000 problems and generated synthetic reasoning traces to train two specialists. Nemotron-3-Nano-CC, with 30 billion total parameters and 3 billion active parameters, received both SFT and RL. Nemotron-3-Ultra-CC, with 550 billion total parameters and 55 billion active parameters, received SFT.
The progression on IOI 2025 makes the value of specialization easy to see. Nano improved from 130 points before post-training to 280 after SFT and 291 after RL. With GenCorrect, our iterative generate-evaluate-refine strategy, it reached 468 points and crossed the gold threshold of 438.3. Ultra-CC reached 502 points with the same test-time strategy.
These experiments also showed that adaptation does not have to look the same at every scale. SFT produced most of Nano's gain, with RL adding a smaller but consistent improvement. For the stronger Ultra model, one SFT epoch was enough to outperform the fully post-trained Nano model across IOI, ICPC, and LiveCodeBench Pro. That finding guided the competition-specific Ultra-CC system used for IOI 2026, which scored 535.4 out of 600.
The IMO project applied the same idea to olympiad mathematics. Starting from Nemotron 3 Ultra, we trained one specialist with SFT and another with RL.
The SFT corpus contained 414,890 quality-filtered examples across 15,818 unique proof problems. It did more than teach final answers. The data covered proof generation, refinement, verification, and meta-verification, so the model learned to construct arguments, identify gaps, respond to critiques, and judge whether a proof was complete. The RL model was trained on 9,597 proof problems selected near the model's capability frontier.
Both post-trained checkpoints outperformed the general-availability model in the development experiments. The SFT checkpoint was strongest in the first search round, while the RL checkpoint achieved the best overall single-checkpoint result. Their strengths were complementary, so the final system used both specialists alongside the general model.
For each IMO problem, the models generated candidate proofs, scored them, produced critiques, and refined the most promising attempts. A separate high-compute stage selected the final submission. The entire system worked in natural language, with no formal prover, external too