REWARD CONSISTENCY: Improving Multi-Objective Alignment from a Data-Centric Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Zhihao, Tong, Yongqi, Zhang, Xin, Zhou, Jun, Wang, Xiting |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Uncovering Safety Risks of Large Language Models through Concept Activation Vector
by: Xu, Zhihao, et al.
Published: (2024)
by: Xu, Zhihao, et al.
Published: (2024)
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment
by: Wang, Sizhe, et al.
Published: (2024)
by: Wang, Sizhe, et al.
Published: (2024)
OrthAlign: Orthogonal Subspace Decomposition for Non-Interfering Multi-Objective Alignment
by: Lin, Liang, et al.
Published: (2025)
by: Lin, Liang, et al.
Published: (2025)
Internal Value Alignment in Large Language Models through Controlled Value Vector Activation
by: Jin, Haoran, et al.
Published: (2025)
by: Jin, Haoran, et al.
Published: (2025)
Preference Orchestrator: Prompt-Aware Multi-Objective Alignment for Large Language Models
by: Liu, Biao, et al.
Published: (2025)
by: Liu, Biao, et al.
Published: (2025)
Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Text
by: Xu, Zhihao, et al.
Published: (2026)
by: Xu, Zhihao, et al.
Published: (2026)
Challenges and Future Directions of Data-Centric AI Alignment
by: Yeh, Min-Hsuan, et al.
Published: (2024)
by: Yeh, Min-Hsuan, et al.
Published: (2024)
Select, Read, and Write: A Multi-Agent Framework of Full-Text-based Related Work Generation
by: Liu, Xiaochuan, et al.
Published: (2025)
by: Liu, Xiaochuan, et al.
Published: (2025)
Learning to Optimize Multi-Objective Alignment Through Dynamic Reward Weighting
by: Lu, Yining, et al.
Published: (2025)
by: Lu, Yining, et al.
Published: (2025)
MOA: Multi-Objective Alignment for Role-Playing Agents
by: Liao, Chonghua, et al.
Published: (2025)
by: Liao, Chonghua, et al.
Published: (2025)
Uncovering Cross-Objective Interference in Multi-Objective Alignment
by: Lu, Yining, et al.
Published: (2026)
by: Lu, Yining, et al.
Published: (2026)
UC-MOA: Utility-Conditioned Multi-Objective Alignment for Distributional Pareto-Optimality
by: Cheng, Zelei, et al.
Published: (2025)
by: Cheng, Zelei, et al.
Published: (2025)
Self-Pluralising Culture Alignment for Large Language Models
by: Xu, Shaoyang, et al.
Published: (2024)
by: Xu, Shaoyang, et al.
Published: (2024)
MOSAIC: Multi-Objective Slice-Aware Iterative Curation for Alignment
by: Dou, Yipu, et al.
Published: (2026)
by: Dou, Yipu, et al.
Published: (2026)
Enhancing Safety of Large Language Models via Embedding Space Separation
by: Zhao, Xu, et al.
Published: (2026)
by: Zhao, Xu, et al.
Published: (2026)
The Re-Label Method For Data-Centric Machine Learning
by: Guo, Tong
Published: (2023)
by: Guo, Tong
Published: (2023)
Data Distribution Matters: A Data-Centric Perspective on Context Compression for Large Language Model
by: Lv, Kangtao, et al.
Published: (2026)
by: Lv, Kangtao, et al.
Published: (2026)
Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards
by: Wang, Haoxiang, et al.
Published: (2024)
by: Wang, Haoxiang, et al.
Published: (2024)
Entropy-based Exploration Conduction for Multi-step Reasoning
by: Zhang, Jinghan, et al.
Published: (2025)
by: Zhang, Jinghan, et al.
Published: (2025)
Alignment Tuning for Large Language Models: A Data-Centric Lens on Alignment Data Pipelines
by: Song, Hwanjun
Published: (2026)
by: Song, Hwanjun
Published: (2026)
AIPO: Improving Training Objective for Iterative Preference Optimization
by: Shen, Yaojie, et al.
Published: (2024)
by: Shen, Yaojie, et al.
Published: (2024)
Improving LLM Safety Alignment with Dual-Objective Optimization
by: Zhao, Xuandong, et al.
Published: (2025)
by: Zhao, Xuandong, et al.
Published: (2025)
Robust Multi-Objective Preference Alignment with Online DPO
by: Gupta, Raghav, et al.
Published: (2025)
by: Gupta, Raghav, et al.
Published: (2025)
Multi-Objective Alignment of Language Models for Personalized Psychotherapy
by: Beikzadeh, Mehrab, et al.
Published: (2026)
by: Beikzadeh, Mehrab, et al.
Published: (2026)
Improving Multi-lingual Alignment Through Soft Contrastive Learning
by: Park, Minsu, et al.
Published: (2024)
by: Park, Minsu, et al.
Published: (2024)
UniARM: Towards a Unified Autoregressive Reward Model for Multi-Objective Test-Time Alignment
by: Xie, Hongyan, et al.
Published: (2026)
by: Xie, Hongyan, et al.
Published: (2026)
MetaAligner: Towards Generalizable Multi-Objective Alignment of Language Models
by: Yang, Kailai, et al.
Published: (2024)
by: Yang, Kailai, et al.
Published: (2024)
DataSculpt: Crafting Data Landscapes for Long-Context LLMs through Multi-Objective Partitioning
by: Lu, Keer, et al.
Published: (2024)
by: Lu, Keer, et al.
Published: (2024)
Pareto Multi-Objective Alignment for Language Models
by: He, Qiang, et al.
Published: (2025)
by: He, Qiang, et al.
Published: (2025)
Understanding and Addressing the Under-Translation Problem from the Perspective of Decoding Objective
by: Shao, Chenze, et al.
Published: (2024)
by: Shao, Chenze, et al.
Published: (2024)
Self-Improvement Towards Pareto Optimality: Mitigating Preference Conflicts in Multi-Objective Alignment
by: Li, Moxin, et al.
Published: (2025)
by: Li, Moxin, et al.
Published: (2025)
Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
by: Guo, Yiju, et al.
Published: (2024)
by: Guo, Yiju, et al.
Published: (2024)
Prototypical Reward Network for Data-Efficient RLHF
by: Zhang, Jinghan, et al.
Published: (2024)
by: Zhang, Jinghan, et al.
Published: (2024)
Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
by: Wang, Haoxiang, et al.
Published: (2024)
by: Wang, Haoxiang, et al.
Published: (2024)
Rendering Data Unlearnable by Exploiting LLM Alignment Mechanisms
by: Zhang, Ruihan, et al.
Published: (2026)
by: Zhang, Ruihan, et al.
Published: (2026)
Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models
by: Li, Chengao, et al.
Published: (2025)
by: Li, Chengao, et al.
Published: (2025)
Prompting Large Language Models for Counterfactual Generation: An Empirical Study
by: Li, Yongqi, et al.
Published: (2023)
by: Li, Yongqi, et al.
Published: (2023)
Towards Next-Generation LLM Training: From the Data-Centric Perspective
by: Liang, Hao, et al.
Published: (2026)
by: Liang, Hao, et al.
Published: (2026)
Unlocking Decoding-time Controllability: Gradient-Free Multi-Objective Alignment with Contrastive Prompts
by: Fu, Tingchen, et al.
Published: (2024)
by: Fu, Tingchen, et al.
Published: (2024)
Adapting Pretrained Language Models for Citation Classification via Self-Supervised Contrastive Learning
by: Li, Tong, et al.
Published: (2025)
by: Li, Tong, et al.
Published: (2025)
Similar Items
-
Uncovering Safety Risks of Large Language Models through Concept Activation Vector
by: Xu, Zhihao, et al.
Published: (2024) -
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment
by: Wang, Sizhe, et al.
Published: (2024) -
OrthAlign: Orthogonal Subspace Decomposition for Non-Interfering Multi-Objective Alignment
by: Lin, Liang, et al.
Published: (2025) -
Internal Value Alignment in Large Language Models through Controlled Value Vector Activation
by: Jin, Haoran, et al.
Published: (2025) -
Preference Orchestrator: Prompt-Aware Multi-Objective Alignment for Large Language Models
by: Liu, Biao, et al.
Published: (2025)