Mode-Conditioning Unlocks Superior Test-Time Scaling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Chen Henry, Goyal, Sachin, Raghunathan, Aditi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909937100128256
author Wu, Chen Henry
Goyal, Sachin
Raghunathan, Aditi
author_facet Wu, Chen Henry
Goyal, Sachin
Raghunathan, Aditi
contents Parallel sampling promises substantial gains in test-time scaling, but its effectiveness is sharply limited by diversity collapse, where models concentrate on a few modes and repeated samples produce the same mistakes. We propose the mode-conditioning (ModC) framework, which explicitly allocates test-time compute across reasoning modes using either specialist models or mode-specific prefixes. ModC consistently improves scaling across controlled graph-search tasks and large-scale reasoning benchmarks, spanning model families and sizes from 0.5B to 7B. On OpenThoughts, fine-tuning Qwen2.5-7B with ModC achieves a 4x efficiency gain over standard training while also improving the maximum attainable Pass@k. We further show that gradient clustering enables ModC without explicit mode labels, yielding up to 10% gains on datasets such as NuminaMath. Finally, we show that ModC improves reinforcement learning (RL) and can further boost diversity-inducing RL methods. These results demonstrate that standard training underutilizes the diversity in data, and that ModC provides a simple, effective remedy for unlocking the full benefits of diversity in test-time scaling.
format Preprint
id arxiv_https___arxiv_org_abs_2512_01127
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mode-Conditioning Unlocks Superior Test-Time Scaling
Wu, Chen Henry
Goyal, Sachin
Raghunathan, Aditi
Machine Learning
Artificial Intelligence
Computation and Language
Parallel sampling promises substantial gains in test-time scaling, but its effectiveness is sharply limited by diversity collapse, where models concentrate on a few modes and repeated samples produce the same mistakes. We propose the mode-conditioning (ModC) framework, which explicitly allocates test-time compute across reasoning modes using either specialist models or mode-specific prefixes. ModC consistently improves scaling across controlled graph-search tasks and large-scale reasoning benchmarks, spanning model families and sizes from 0.5B to 7B. On OpenThoughts, fine-tuning Qwen2.5-7B with ModC achieves a 4x efficiency gain over standard training while also improving the maximum attainable Pass@k. We further show that gradient clustering enables ModC without explicit mode labels, yielding up to 10% gains on datasets such as NuminaMath. Finally, we show that ModC improves reinforcement learning (RL) and can further boost diversity-inducing RL methods. These results demonstrate that standard training underutilizes the diversity in data, and that ModC provides a simple, effective remedy for unlocking the full benefits of diversity in test-time scaling.
title Mode-Conditioning Unlocks Superior Test-Time Scaling
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2512.01127