SPARD: Self-Paced Curriculum for RL Alignment via Integrating Reward Dynamics and Data Utility

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhi, Xuyang, zhou, Peilun, Lu, Chengqiang, Lv, Hang, Liang, Yiwei, Zhang, Rongyang, Gao, Yan, WU, YI, Hu, Yao, Gu, Hongchao, Lian, Defu, Wang, Hao, Chen, Enhong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908949223047168
author Zhi, Xuyang
zhou, Peilun
Lu, Chengqiang
Lv, Hang
Liang, Yiwei
Zhang, Rongyang
Gao, Yan
WU, YI
Hu, Yao
Gu, Hongchao
Lian, Defu
Wang, Hao
Chen, Enhong
author_facet Zhi, Xuyang
zhou, Peilun
Lu, Chengqiang
Lv, Hang
Liang, Yiwei
Zhang, Rongyang
Gao, Yan
WU, YI
Hu, Yao
Gu, Hongchao
Lian, Defu
Wang, Hao
Chen, Enhong
contents The evolution of Large Language Models (LLMs) is shifting the focus from single, verifiable tasks toward complex, open-ended real-world scenarios, imposing significant challenges on the post-training phase. In these settings, the scale and complexity of reward systems have grown significantly, transitioning toward multi-objective formulations that encompass a comprehensive spectrum of model capabilities and application contexts. However, traditional methods typically rely on fixed reward weights, ignoring non-stationary learning dynamics and struggling with data heterogeneity across dimensions. To address these issues, we propose SPARD, a framework that establishes an automated, self-paced curriculum by perceiving learning progress to dynamically adjust multi-objective reward weights and data importance, thereby synchronizing learning intent with data utility for optimal performance. Extensive experiments across multiple benchmarks demonstrate that SPARD significantly enhances model capabilities across all domains.
format Preprint
id arxiv_https___arxiv_org_abs_2604_07837
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SPARD: Self-Paced Curriculum for RL Alignment via Integrating Reward Dynamics and Data Utility
Zhi, Xuyang
zhou, Peilun
Lu, Chengqiang
Lv, Hang
Liang, Yiwei
Zhang, Rongyang
Gao, Yan
WU, YI
Hu, Yao
Gu, Hongchao
Lian, Defu
Wang, Hao
Chen, Enhong
Artificial Intelligence
The evolution of Large Language Models (LLMs) is shifting the focus from single, verifiable tasks toward complex, open-ended real-world scenarios, imposing significant challenges on the post-training phase. In these settings, the scale and complexity of reward systems have grown significantly, transitioning toward multi-objective formulations that encompass a comprehensive spectrum of model capabilities and application contexts. However, traditional methods typically rely on fixed reward weights, ignoring non-stationary learning dynamics and struggling with data heterogeneity across dimensions. To address these issues, we propose SPARD, a framework that establishes an automated, self-paced curriculum by perceiving learning progress to dynamically adjust multi-objective reward weights and data importance, thereby synchronizing learning intent with data utility for optimal performance. Extensive experiments across multiple benchmarks demonstrate that SPARD significantly enhances model capabilities across all domains.
title SPARD: Self-Paced Curriculum for RL Alignment via Integrating Reward Dynamics and Data Utility
topic Artificial Intelligence
url https://arxiv.org/abs/2604.07837