DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhou, Yang, Jin, Can, Dong, Zihan, Wang, Zhepeng, Yang, Yanting, Zhao, Shiyu, Li, Lei, Bao, Runxue, Xie, Yaochen, Metaxas, Dimitris N.
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913107527335936
author Zhou, Yang
Jin, Can
Dong, Zihan
Wang, Zhepeng
Yang, Yanting
Zhao, Shiyu
Li, Lei
Bao, Runxue
Xie, Yaochen
Metaxas, Dimitris N.
author_facet Zhou, Yang
Jin, Can
Dong, Zihan
Wang, Zhepeng
Yang, Yanting
Zhao, Shiyu
Li, Lei
Bao, Runxue
Xie, Yaochen
Metaxas, Dimitris N.
contents Reinforcement learning improves the reasoning ability of large language models but remains costly and sample-inefficient, as many rollouts provide weak learning signals. Difficulty-aware data selection methods attempt to address this by prioritizing moderately difficult prompts, yet our analysis reveals three limitations: difficulty estimates become inaccurate under policy drift, data selection alone yields limited final-performance gains, and inference efficiency remains largely unchanged. These findings suggest that efficient and effective RL requires more than filtering by difficulty: the policy should learn to solve hard tasks while producing concise responses for easy ones. To this end, we propose **Dare**, a unified framework that co-evolves difficulty estimation with the policy via self-normalized importance sampling, maintains diverse difficulty coverage through a symmetric Beta sampling distribution, and applies tailored training strategies across difficulty tiers with adaptive compute allocation. Extensive experiments across multiple models and domains demonstrate that **Dare** consistently outperforms existing methods in training efficiency, final effectiveness, and inference efficiency, producing more concise responses on easy tasks while improving correctness on hard ones. Code is available at https://github.com/EtaYang10th/DARE.
format Preprint
id arxiv_https___arxiv_org_abs_2605_09188
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation
Zhou, Yang
Jin, Can
Dong, Zihan
Wang, Zhepeng
Yang, Yanting
Zhao, Shiyu
Li, Lei
Bao, Runxue
Xie, Yaochen
Metaxas, Dimitris N.
Machine Learning
Artificial Intelligence
Reinforcement learning improves the reasoning ability of large language models but remains costly and sample-inefficient, as many rollouts provide weak learning signals. Difficulty-aware data selection methods attempt to address this by prioritizing moderately difficult prompts, yet our analysis reveals three limitations: difficulty estimates become inaccurate under policy drift, data selection alone yields limited final-performance gains, and inference efficiency remains largely unchanged. These findings suggest that efficient and effective RL requires more than filtering by difficulty: the policy should learn to solve hard tasks while producing concise responses for easy ones. To this end, we propose **Dare**, a unified framework that co-evolves difficulty estimation with the policy via self-normalized importance sampling, maintains diverse difficulty coverage through a symmetric Beta sampling distribution, and applies tailored training strategies across difficulty tiers with adaptive compute allocation. Extensive experiments across multiple models and domains demonstrate that **Dare** consistently outperforms existing methods in training efficiency, final effectiveness, and inference efficiency, producing more concise responses on easy tasks while improving correctness on hard ones. Code is available at https://github.com/EtaYang10th/DARE.
title DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.09188