SPARD: Self-Paced Curriculum for RL Alignment via Integrating Reward Dynamics and Data Utility
Fuente:
arXiv
Saved in:
| Main Authors: | Zhi, Xuyang, zhou, Peilun, Lu, Chengqiang, Lv, Hang, Liang, Yiwei, Zhang, Rongyang, Gao, Yan, WU, YI, Hu, Yao, Gu, Hongchao, Lian, Defu, Wang, Hao, Chen, Enhong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SelfAug: Mitigating Catastrophic Forgetting in Retrieval-Augmented Generation via Distribution Self-Alignment
by: Huang, Yuqing, et al.
Published: (2025)
by: Huang, Yuqing, et al.
Published: (2025)
IE as Cache: Information Extraction Enhanced Agentic Reasoning
by: Lv, Hang, et al.
Published: (2026)
by: Lv, Hang, et al.
Published: (2026)
Learning from Emptiness: De-biasing Listwise Rerankers with Content-Agnostic Probability Calibration
by: Lv, Hang, et al.
Published: (2026)
by: Lv, Hang, et al.
Published: (2026)
SpecSteer: Synergizing Local Context and Global Reasoning for Efficient Personalized Generation
by: Lv, Hang, et al.
Published: (2026)
by: Lv, Hang, et al.
Published: (2026)
CoSteer: Collaborative Decoding-Time Personalization via Local Delta Steering
by: Lv, Hang, et al.
Published: (2025)
by: Lv, Hang, et al.
Published: (2025)
RAPID: Efficient Retrieval-Augmented Long Text Generation with Writing Planning and Information Discovery
by: Gu, Hongchao, et al.
Published: (2025)
by: Gu, Hongchao, et al.
Published: (2025)
Is Softmax Loss All You Need? A Principled Analysis of Softmax-family Loss
by: Pu, Yuanhao, et al.
Published: (2026)
by: Pu, Yuanhao, et al.
Published: (2026)
Beyond Surrogates: A Quantitative Analysis for Inter-Metric Relationships
by: Pu, Yuanhao, et al.
Published: (2026)
by: Pu, Yuanhao, et al.
Published: (2026)
Efficient Machine Unlearning via Influence Approximation
by: Liu, Jiawei, et al.
Published: (2025)
by: Liu, Jiawei, et al.
Published: (2025)
LLM Cache Bandit Revisited: Addressing Query Heterogeneity for Cost-Effective LLM Inference
by: Yang, Hantao, et al.
Published: (2025)
by: Yang, Hantao, et al.
Published: (2025)
Securing Recommender System via Cooperative Training
by: Wang, Qingyang, et al.
Published: (2024)
by: Wang, Qingyang, et al.
Published: (2024)
Cosmology of Infinite Self-Reference K: Coupling Infinite Nested Self-Referential Reconstruction with Existing Particle Theories
by: zhou, changzheng, et al.
Published: (2025)
by: zhou, changzheng, et al.
Published: (2025)
Distributionally Robust Self Paced Curriculum Reinforcement Learning
by: Satheesh, Anirudh, et al.
Published: (2025)
by: Satheesh, Anirudh, et al.
Published: (2025)
Diagnosing and Mitigating System Bias in Self-Rewarding RL
by: Tan, Chuyi, et al.
Published: (2025)
by: Tan, Chuyi, et al.
Published: (2025)
A novel fault diagnosis method for imbalanced datasets based on MCNN‐Transformer model in industrial processes
by: Rongyang Lu
Published: (2024)
by: Rongyang Lu
Published: (2024)
Analytical and Empirical Study of Herding Effects in Recommendation Systems
by: Xie, Hong, et al.
Published: (2024)
by: Xie, Hong, et al.
Published: (2024)
Multi-agent Multi-armed Bandits with Stochastic Sharable Arm Capacities
by: Xie, Hong, et al.
Published: (2024)
by: Xie, Hong, et al.
Published: (2024)
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
by: Huang, Yu, et al.
Published: (2026)
by: Huang, Yu, et al.
Published: (2026)
Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners
by: Muslimani, Calarina, et al.
Published: (2025)
by: Muslimani, Calarina, et al.
Published: (2025)
∞-Groupoid Cosmology E; Emergence of Spacetime and Matter from the Self-Referential Structure of ∞-Groupoids
by: zhou, Changzheng, et al.
Published: (2026)
by: zhou, Changzheng, et al.
Published: (2026)
EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling
by: Luo, Xin, et al.
Published: (2025)
by: Luo, Xin, et al.
Published: (2025)
UniMEL: A Unified Framework for Multimodal Entity Linking with Large Language Models
by: Qi, Liu, et al.
Published: (2024)
by: Qi, Liu, et al.
Published: (2024)
RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering
by: Zhang, Rongyang, et al.
Published: (2025)
by: Zhang, Rongyang, et al.
Published: (2025)
Prompting is not Enough: Exploring Knowledge Integration and Controllable Generation
by: Shen, Tingjia, et al.
Published: (2025)
by: Shen, Tingjia, et al.
Published: (2025)
The Ultimate Number Domain of Complex Numbers: Fieldoid F; Higher-Order Categories and the Structural Ontology of the Self-Referential Universe
by: zhou, changzheng, et al.
Published: (2026)
by: zhou, changzheng, et al.
Published: (2026)
ChemEval: A Comprehensive Multi-Level Chemical Evaluation for Large Language Models
by: Huang, Yuqing, et al.
Published: (2024)
by: Huang, Yuqing, et al.
Published: (2024)
Exploring User Retrieval Integration towards Large Language Models for Cross-Domain Sequential Recommendation
by: Shen, Tingjia, et al.
Published: (2024)
by: Shen, Tingjia, et al.
Published: (2024)
Using Aristotle API for AI-Assisted Theorem Proving in Lean 4: A Formalisation Case Study of the Grasshopper Problem
by: Lau, Gabriel Rongyang
Published: (2026)
by: Lau, Gabriel Rongyang
Published: (2026)
Why Thinking Hurts: Diagnosing and Rectifying Linguistic Inertia in Large Language Models for Recommendation
by: Zhang, Luankang, et al.
Published: (2026)
by: Zhang, Luankang, et al.
Published: (2026)
PiXTime: A Model for Federated Time Series Forecasting with Heterogeneous Data across Nodes
by: Zhou, Yiming, et al.
Published: (2026)
by: Zhou, Yiming, et al.
Published: (2026)
Learning to Substitute Components for Compositional Generalization
by: Li, Zhaoyi, et al.
Published: (2025)
by: Li, Zhaoyi, et al.
Published: (2025)
Breaking Determinism: Fuzzy Modeling of Sequential Recommendation Using Discrete State Space Diffusion Model
by: Xie, Wenjia, et al.
Published: (2024)
by: Xie, Wenjia, et al.
Published: (2024)
NDCG-Consistent Softmax Approximation with Accelerated Convergence
by: Pu, Yuanhao, et al.
Published: (2025)
by: Pu, Yuanhao, et al.
Published: (2025)
Model Specific Task Similarity for Vision Language Model Selection via Layer Conductance
by: Yang, Wei, et al.
Published: (2026)
by: Yang, Wei, et al.
Published: (2026)
Adaptive Sampled Softmax with Inverted Multi-Index: Methods, Theory and Applications
by: Chen, Jin, et al.
Published: (2025)
by: Chen, Jin, et al.
Published: (2025)
Learning Deep Tree-based Retriever for Efficient Recommendation: Theory and Method
by: Liu, Ze, et al.
Published: (2024)
by: Liu, Ze, et al.
Published: (2024)
MDAP: A Multi-view Disentangled and Adaptive Preference Learning Framework for Cross-Domain Recommendation
by: Tong, Junxiong, et al.
Published: (2024)
by: Tong, Junxiong, et al.
Published: (2024)
Model Selection for Average Reward RL with Application to Utility Maximization in Repeated Games
by: Masoumian, Alireza, et al.
Published: (2024)
by: Masoumian, Alireza, et al.
Published: (2024)
Demystifying Design Choices of Reinforcement Fine-tuning: A Batched Contextual Bandit Learning Perspective
by: Xie, Hong, et al.
Published: (2026)
by: Xie, Hong, et al.
Published: (2026)
SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection
by: Chen, Shuhao, et al.
Published: (2026)
by: Chen, Shuhao, et al.
Published: (2026)
Similar Items
-
SelfAug: Mitigating Catastrophic Forgetting in Retrieval-Augmented Generation via Distribution Self-Alignment
by: Huang, Yuqing, et al.
Published: (2025) -
IE as Cache: Information Extraction Enhanced Agentic Reasoning
by: Lv, Hang, et al.
Published: (2026) -
Learning from Emptiness: De-biasing Listwise Rerankers with Content-Agnostic Probability Calibration
by: Lv, Hang, et al.
Published: (2026) -
SpecSteer: Synergizing Local Context and Global Reasoning for Efficient Personalized Generation
by: Lv, Hang, et al.
Published: (2026) -
CoSteer: Collaborative Decoding-Time Personalization via Local Delta Steering
by: Lv, Hang, et al.
Published: (2025)