Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Boxin, Lee, Chankyu, Lee, Nayeon, Lin, Sheng-Chieh, Dai, Wenliang, Chen, Yang, Chen, Yangyi, Yang, Zhuolin, Liu, Zihan, Shoeybi, Mohammad, Catanzaro, Bryan, Ping, Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911547734884352
author Wang, Boxin
Lee, Chankyu
Lee, Nayeon
Lin, Sheng-Chieh
Dai, Wenliang
Chen, Yang
Chen, Yangyi
Yang, Zhuolin
Liu, Zihan
Shoeybi, Mohammad
Catanzaro, Bryan
Ping, Wei
author_facet Wang, Boxin
Lee, Chankyu
Lee, Nayeon
Lin, Sheng-Chieh
Dai, Wenliang
Chen, Yang
Chen, Yangyi
Yang, Zhuolin
Liu, Zihan
Shoeybi, Mohammad
Catanzaro, Bryan
Ping, Wei
contents Building general-purpose reasoning models with reinforcement learning (RL) entails substantial cross-domain heterogeneity, including large variation in inference-time response lengths and verification latency. Such variability complicates the RL infrastructure, slows training, and makes training curriculum (e.g., response length extension) and hyperparameter selection challenging. In this work, we propose cascaded domain-wise reinforcement learning (Cascade RL) to develop Nemotron-Cascade, capable of operating in both instruct and deep thinking modes, without any performance gap relative to a thinking-only counterpart. Departing from conventional approaches that blend heterogeneous prompts from different domains, Cascade RL orchestrates sequential, domain-wise RL, reducing engineering complexity and delivering state-of-the-art performance across a wide range of benchmarks. Notably, RLHF for alignment, when used as a pre-step, boosts the model's reasoning ability far beyond mere preference optimization, and subsequent domain-wise RLVR stages rarely degrade the benchmark performance attained in earlier domains and may even improve it (see an illustration in Figure 1). Our 14B model, after RL, outperforms its SFT teacher, DeepSeek-R1-0528, on LiveCodeBench v5/v6/Pro and achieves silver-medal performance in the 2025 International Olympiad in Informatics (IOI). We transparently share our training and data recipes.
format Preprint
id arxiv_https___arxiv_org_abs_2512_13607
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
Wang, Boxin
Lee, Chankyu
Lee, Nayeon
Lin, Sheng-Chieh
Dai, Wenliang
Chen, Yang
Chen, Yangyi
Yang, Zhuolin
Liu, Zihan
Shoeybi, Mohammad
Catanzaro, Bryan
Ping, Wei
Computation and Language
Artificial Intelligence
Machine Learning
Building general-purpose reasoning models with reinforcement learning (RL) entails substantial cross-domain heterogeneity, including large variation in inference-time response lengths and verification latency. Such variability complicates the RL infrastructure, slows training, and makes training curriculum (e.g., response length extension) and hyperparameter selection challenging. In this work, we propose cascaded domain-wise reinforcement learning (Cascade RL) to develop Nemotron-Cascade, capable of operating in both instruct and deep thinking modes, without any performance gap relative to a thinking-only counterpart. Departing from conventional approaches that blend heterogeneous prompts from different domains, Cascade RL orchestrates sequential, domain-wise RL, reducing engineering complexity and delivering state-of-the-art performance across a wide range of benchmarks. Notably, RLHF for alignment, when used as a pre-step, boosts the model's reasoning ability far beyond mere preference optimization, and subsequent domain-wise RLVR stages rarely degrade the benchmark performance attained in earlier domains and may even improve it (see an illustration in Figure 1). Our 14B model, after RL, outperforms its SFT teacher, DeepSeek-R1-0528, on LiveCodeBench v5/v6/Pro and achieves silver-medal performance in the 2025 International Olympiad in Informatics (IOI). We transparently share our training and data recipes.
title Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2512.13607