Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Du, Wei, Toshniwal, Shubham, Kisacanin, Branislav, Mahdavi, Sadegh, Moshkov, Ivan, Armstrong, George, Ge, Stephen, Minasyan, Edgar, Chen, Feng, Gitman, Igor
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908718312980480
author Du, Wei
Toshniwal, Shubham
Kisacanin, Branislav
Mahdavi, Sadegh
Moshkov, Ivan
Armstrong, George
Ge, Stephen
Minasyan, Edgar
Chen, Feng
Gitman, Igor
author_facet Du, Wei
Toshniwal, Shubham
Kisacanin, Branislav
Mahdavi, Sadegh
Moshkov, Ivan
Armstrong, George
Ge, Stephen
Minasyan, Edgar
Chen, Feng
Gitman, Igor
contents High-quality mathematical reasoning supervision requires diverse reasoning styles, long-form traces, and effective tool integration, capabilities that existing datasets provide only in limited form. Leveraging the multi-mode generation ability of gpt-oss-120b, we introduce Nemotron-Math, a large-scale mathematical reasoning dataset containing 7.5M solution traces across high, medium, and low reasoning modes, each available both with and without Python tool-integrated reasoning (TIR). The dataset integrates 85K curated AoPS problems with 262K community-sourced StackExchange-Math problems, combining structured competition tasks with diverse real-world mathematical queries. We conduct controlled evaluations to assess the dataset quality. Nemotron-Math consistently outperforms the original OpenMathReasoning on matched AoPS problems. Incorporating StackExchange-Math substantially improves robustness and generalization, especially on HLE-Math, while preserving accuracy on math competition benchmarks. To support efficient long-context training, we develop a sequential bucketed strategy that accelerates 128K context-length fine-tuning by 2--3$\times$ without significant accuracy loss. Overall, Nemotron-Math enables state-of-the-art performance, including 100\% maj@16 accuracy on AIME 2024 and 2025 with Python TIR.
format Preprint
id arxiv_https___arxiv_org_abs_2512_15489
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision
Du, Wei
Toshniwal, Shubham
Kisacanin, Branislav
Mahdavi, Sadegh
Moshkov, Ivan
Armstrong, George
Ge, Stephen
Minasyan, Edgar
Chen, Feng
Gitman, Igor
Artificial Intelligence
High-quality mathematical reasoning supervision requires diverse reasoning styles, long-form traces, and effective tool integration, capabilities that existing datasets provide only in limited form. Leveraging the multi-mode generation ability of gpt-oss-120b, we introduce Nemotron-Math, a large-scale mathematical reasoning dataset containing 7.5M solution traces across high, medium, and low reasoning modes, each available both with and without Python tool-integrated reasoning (TIR). The dataset integrates 85K curated AoPS problems with 262K community-sourced StackExchange-Math problems, combining structured competition tasks with diverse real-world mathematical queries. We conduct controlled evaluations to assess the dataset quality. Nemotron-Math consistently outperforms the original OpenMathReasoning on matched AoPS problems. Incorporating StackExchange-Math substantially improves robustness and generalization, especially on HLE-Math, while preserving accuracy on math competition benchmarks. To support efficient long-context training, we develop a sequential bucketed strategy that accelerates 128K context-length fine-tuning by 2--3$\times$ without significant accuracy loss. Overall, Nemotron-Math enables state-of-the-art performance, including 100\% maj@16 accuracy on AIME 2024 and 2025 with Python TIR.
title Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision
topic Artificial Intelligence
url https://arxiv.org/abs/2512.15489