DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Yuhang, Zhu, Jing, Qian, Shengyi, Zhao, Zhuokai, Wang, Xiyao, Liu, Xiaoyu, Li, Ming, Xu, Paiheng, Ai, Wei, Huang, Furong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918146824208384
author Zhou, Yuhang
Zhu, Jing
Qian, Shengyi
Zhao, Zhuokai
Wang, Xiyao
Liu, Xiaoyu
Li, Ming
Xu, Paiheng
Ai, Wei
Huang, Furong
author_facet Zhou, Yuhang
Zhu, Jing
Qian, Shengyi
Zhao, Zhuokai
Wang, Xiyao
Liu, Xiaoyu
Li, Ming
Xu, Paiheng
Ai, Wei
Huang, Furong
contents Large Language Models (LLMs) are increasingly aligned with human preferences through Reinforcement Learning from Human Feedback (RLHF). Among RLHF methods, Group Relative Policy Optimization (GRPO) has gained attention for its simplicity and strong performance, notably eliminating the need for a learned value function. However, GRPO implicitly assumes a balanced domain distribution and uniform semantic alignment across groups, assumptions that rarely hold in real-world datasets. When applied to multi-domain, imbalanced data, GRPO disproportionately optimizes for dominant domains, neglecting underrepresented ones and resulting in poor generalization and fairness. We propose Domain-Informed Self-Consistency Policy Optimization (DISCO), a principled extension to GRPO that addresses inter-group imbalance with two key innovations. Domain-aware reward scaling counteracts frequency bias by reweighting optimization based on domain prevalence. Difficulty-aware reward scaling leverages prompt-level self-consistency to identify and prioritize uncertain prompts that offer greater learning value. Together, these strategies promote more equitable and effective policy learning across domains. Extensive experiments across multiple LLMs and skewed training distributions show that DISCO improves generalization, outperforms existing GRPO variants by 5% on Qwen3 models, and sets new state-of-the-art results on multi-domain alignment benchmarks. Our code and data are available at https://github.com/Tonyzhou98/disco_grpo.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15074
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data
Zhou, Yuhang
Zhu, Jing
Qian, Shengyi
Zhao, Zhuokai
Wang, Xiyao
Liu, Xiaoyu
Li, Ming
Xu, Paiheng
Ai, Wei
Huang, Furong
Computation and Language
Artificial Intelligence
Machine Learning
Large Language Models (LLMs) are increasingly aligned with human preferences through Reinforcement Learning from Human Feedback (RLHF). Among RLHF methods, Group Relative Policy Optimization (GRPO) has gained attention for its simplicity and strong performance, notably eliminating the need for a learned value function. However, GRPO implicitly assumes a balanced domain distribution and uniform semantic alignment across groups, assumptions that rarely hold in real-world datasets. When applied to multi-domain, imbalanced data, GRPO disproportionately optimizes for dominant domains, neglecting underrepresented ones and resulting in poor generalization and fairness. We propose Domain-Informed Self-Consistency Policy Optimization (DISCO), a principled extension to GRPO that addresses inter-group imbalance with two key innovations. Domain-aware reward scaling counteracts frequency bias by reweighting optimization based on domain prevalence. Difficulty-aware reward scaling leverages prompt-level self-consistency to identify and prioritize uncertain prompts that offer greater learning value. Together, these strategies promote more equitable and effective policy learning across domains. Extensive experiments across multiple LLMs and skewed training distributions show that DISCO improves generalization, outperforms existing GRPO variants by 5% on Qwen3 models, and sets new state-of-the-art results on multi-domain alignment benchmarks. Our code and data are available at https://github.com/Tonyzhou98/disco_grpo.
title DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.15074