HyperAIHyperAI

Command Palette

Search for a command to run...

Fine-Tuned Nemotron Achieves Gold-Level Scores at IOI and IMO

Nemotron Labs has achieved gold-medal performance at both the 2026 International Olympiad in Informatics and the 2026 International Mathematical Olympiad, establishing a reproducible framework for transforming general-purpose foundation models into elite domain specialists. By leveraging a unified fine-tuning and inference strategy, the team demonstrated that competitive excellence at the frontier of algorithmic programming and formal mathematics can be systematically engineered rather than achieved through brute-force compute alone. The methodology centers on a four-step specialization recipe. Teams began with the Nemotron 3 base model, curated high-quality domain datasets, applied standard post-training techniques including supervised fine-tuning and reinforcement learning, and integrated a feedback-driven inference loop that generates, evaluates, and refines candidate solutions. This approach eliminates the need to train custom foundational architectures for each challenge, proving that targeted adaptation yields superior results across diverse cognitive tasks. For the IOI competition, the focus was on algorithmic problem-solving under strict time and internet-access constraints. The Nemotron-3-Ultra-CC model, a 550-billion-parameter variant with 55 billion active parameters, underwent supervised fine-tuning on 22,000 curated programming problems and synthetic reasoning traces. During a live, unsupervised benchmark run matching official competition conditions, the model achieved a score of 535.4 out of 600, significantly surpassing the gold threshold of 361.12 and outperforming the top human result. The iterative GenCorrect inference strategy played a critical role, enabling the model to progressively refine code submissions across multiple evaluation rounds. The IMO project applied the same architectural philosophy to rigorous natural-language mathematical proofs. Starting from Nemotron 3 Ultra, specialists were trained on a corpus of over 414,000 filtered proof examples, covering generation, verification, and critique identification. A reinforcement learning checkpoint was trained on near-capability-frontier problems. Rather than relying on a single model, the system combined complementary supervised and reinforcement learning checkpoints within a generate-verify-refine pipeline. Operating entirely in natural language without external tools or formal provers, the system earned 30 out of 42 points, exceeding the official gold threshold of 29 and securing full credit on four of six problems. The results underscore a critical shift in artificial intelligence engineering: optimal performance emerges from co-designing model architectures, training data, and test-time inference workflows. Neither fine-tuning alone nor massive sampling proved sufficient; success required specialized checkpoints paired with transparent, iterative refinement loops. The team has since released the full collection of checkpoints, training datasets, evaluation benchmarks, and inference pipelines on Hugging Face and the NeMo-Skills repository. This open release invites the broader developer community to build upon a validated recipe for creating high-performance, competition-grade AI specialists.

Related Links