DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866913107527335936 |
|---|---|
| author | Zhou, Yang Jin, Can Dong, Zihan Wang, Zhepeng Yang, Yanting Zhao, Shiyu Li, Lei Bao, Runxue Xie, Yaochen Metaxas, Dimitris N. |
| author_facet | Zhou, Yang Jin, Can Dong, Zihan Wang, Zhepeng Yang, Yanting Zhao, Shiyu Li, Lei Bao, Runxue Xie, Yaochen Metaxas, Dimitris N. |
| contents | Reinforcement learning improves the reasoning ability of large language models but remains costly and sample-inefficient, as many rollouts provide weak learning signals. Difficulty-aware data selection methods attempt to address this by prioritizing moderately difficult prompts, yet our analysis reveals three limitations: difficulty estimates become inaccurate under policy drift, data selection alone yields limited final-performance gains, and inference efficiency remains largely unchanged. These findings suggest that efficient and effective RL requires more than filtering by difficulty: the policy should learn to solve hard tasks while producing concise responses for easy ones. To this end, we propose **Dare**, a unified framework that co-evolves difficulty estimation with the policy via self-normalized importance sampling, maintains diverse difficulty coverage through a symmetric Beta sampling distribution, and applies tailored training strategies across difficulty tiers with adaptive compute allocation. Extensive experiments across multiple models and domains demonstrate that **Dare** consistently outperforms existing methods in training efficiency, final effectiveness, and inference efficiency, producing more concise responses on easy tasks while improving correctness on hard ones. Code is available at https://github.com/EtaYang10th/DARE. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_09188 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation Zhou, Yang Jin, Can Dong, Zihan Wang, Zhepeng Yang, Yanting Zhao, Shiyu Li, Lei Bao, Runxue Xie, Yaochen Metaxas, Dimitris N. Machine Learning Artificial Intelligence Reinforcement learning improves the reasoning ability of large language models but remains costly and sample-inefficient, as many rollouts provide weak learning signals. Difficulty-aware data selection methods attempt to address this by prioritizing moderately difficult prompts, yet our analysis reveals three limitations: difficulty estimates become inaccurate under policy drift, data selection alone yields limited final-performance gains, and inference efficiency remains largely unchanged. These findings suggest that efficient and effective RL requires more than filtering by difficulty: the policy should learn to solve hard tasks while producing concise responses for easy ones. To this end, we propose **Dare**, a unified framework that co-evolves difficulty estimation with the policy via self-normalized importance sampling, maintains diverse difficulty coverage through a symmetric Beta sampling distribution, and applies tailored training strategies across difficulty tiers with adaptive compute allocation. Extensive experiments across multiple models and domains demonstrate that **Dare** consistently outperforms existing methods in training efficiency, final effectiveness, and inference efficiency, producing more concise responses on easy tasks while improving correctness on hard ones. Code is available at https://github.com/EtaYang10th/DARE. |
| title | DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2605.09188 |