Pairwise Calibrated Rewards for Pluralistic Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Halpern, Daniel, Micha, Evi, Procaccia, Ariel D., Shapira, Itai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Axioms for AI Alignment from Human Feedback
by: Ge, Luise, et al.
Published: (2024)
by: Ge, Luise, et al.
Published: (2024)
Strategic Classification With Externalities
by: Hossain, Safwan, et al.
Published: (2024)
by: Hossain, Safwan, et al.
Published: (2024)
How RLHF Amplifies Sycophancy
by: Shapira, Itai, et al.
Published: (2026)
by: Shapira, Itai, et al.
Published: (2026)
Clone-Robust AI Alignment
by: Procaccia, Ariel D., et al.
Published: (2025)
by: Procaccia, Ariel D., et al.
Published: (2025)
Generative Social Choice
by: Fish, Sara, et al.
Published: (2023)
by: Fish, Sara, et al.
Published: (2023)
Direct Alignment with Heterogeneous Preferences
by: Shirali, Ali, et al.
Published: (2025)
by: Shirali, Ali, et al.
Published: (2025)
Proportional Fairness in Non-Centroid Clustering
by: Caragiannis, Ioannis, et al.
Published: (2024)
by: Caragiannis, Ioannis, et al.
Published: (2024)
Pluralistic Alignment Over Time
by: Klassen, Toryn Q., et al.
Published: (2024)
by: Klassen, Toryn Q., et al.
Published: (2024)
Robust AI Evaluation through Maximal Lotteries
by: Khalaf, Hadi, et al.
Published: (2026)
by: Khalaf, Hadi, et al.
Published: (2026)
Multi-objective Reinforcement Learning: A Tool for Pluralistic Alignment
by: Vamplew, Peter, et al.
Published: (2024)
by: Vamplew, Peter, et al.
Published: (2024)
Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
by: Guo, Hanze, et al.
Published: (2025)
by: Guo, Hanze, et al.
Published: (2025)
APPA: Adaptive Preference Pluralistic Alignment for Fair Federated RLHF of LLMs
by: Srewa, Mahmoud, et al.
Published: (2026)
by: Srewa, Mahmoud, et al.
Published: (2026)
Learning Social Welfare Functions
by: Pardeshi, Kanad Shrikar, et al.
Published: (2024)
by: Pardeshi, Kanad Shrikar, et al.
Published: (2024)
Pluralistic Alignment for Healthcare: A Role-Driven Framework
by: Zhong, Jiayou, et al.
Published: (2025)
by: Zhong, Jiayou, et al.
Published: (2025)
MixDPO: Modeling Preference Strength for Pluralistic Alignment
by: Imai, Saki, et al.
Published: (2026)
by: Imai, Saki, et al.
Published: (2026)
Generative Social Choice: The Next Generation
by: Boehmer, Niclas, et al.
Published: (2025)
by: Boehmer, Niclas, et al.
Published: (2025)
VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare
by: Shetty, Anudeex, et al.
Published: (2025)
by: Shetty, Anudeex, et al.
Published: (2025)
VISPA: Pluralistic Alignment via Automatic Value Selection and Activation
by: Zheng, Shenyan, et al.
Published: (2026)
by: Zheng, Shenyan, et al.
Published: (2026)
Adaptive Alignment: Dynamic Preference Adjustments via Multi-Objective Reinforcement Learning for Pluralistic AI
by: Harland, Hadassah, et al.
Published: (2024)
by: Harland, Hadassah, et al.
Published: (2024)
Being Considerate as a Pathway Towards Pluralistic Alignment for Agentic AI
by: Alamdari, Parand A., et al.
Published: (2024)
by: Alamdari, Parand A., et al.
Published: (2024)
A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI
by: Karagoz, Atahan
Published: (2026)
by: Karagoz, Atahan
Published: (2026)
Policy Aggregation
by: Alamdari, Parand A., et al.
Published: (2024)
by: Alamdari, Parand A., et al.
Published: (2024)
$i$REPO: $i$mplicit Reward Pairwise Difference based Empirical Preference Optimization
by: Le, Long Tan, et al.
Published: (2024)
by: Le, Long Tan, et al.
Published: (2024)
Beyond Pairwise: Empowering LLM Alignment With Ranked Choice Modeling
by: Tang, Yuxuan, et al.
Published: (2025)
by: Tang, Yuxuan, et al.
Published: (2025)
Decipher the Modality Gap in Multimodal Contrastive Learning: From Convergent Representations to Pairwise Alignment
by: Yi, Lingjie, et al.
Published: (2025)
by: Yi, Lingjie, et al.
Published: (2025)
Alternates, Assemble! Selecting Optimal Alternates for Citizens' Assemblies
by: Assos, Angelos, et al.
Published: (2025)
by: Assos, Angelos, et al.
Published: (2025)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
by: Wang, Chaoqi, et al.
Published: (2025)
by: Wang, Chaoqi, et al.
Published: (2025)
PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling
by: Jian, Ai, et al.
Published: (2025)
by: Jian, Ai, et al.
Published: (2025)
Local Pairwise Distance Matching for Backpropagation-Free Reinforcement Learning
by: Tanneberg, Daniel
Published: (2025)
by: Tanneberg, Daniel
Published: (2025)
SOAP: Improving and Stabilizing Shampoo using Adam
by: Vyas, Nikhil, et al.
Published: (2024)
by: Vyas, Nikhil, et al.
Published: (2024)
CHARM: Calibrating Reward Models With Chatbot Arena Scores
by: Zhu, Xiao, et al.
Published: (2025)
by: Zhu, Xiao, et al.
Published: (2025)
Evaluating Large Language Models for Fair and Reliable Organ Allocation
by: Kim, Brian Hyeongseok, et al.
Published: (2025)
by: Kim, Brian Hyeongseok, et al.
Published: (2025)
Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport
by: Ma, Rachel, et al.
Published: (2026)
by: Ma, Rachel, et al.
Published: (2026)
The Reward Model Selection Crisis in Personalized Alignment
by: Rezk, Fady, et al.
Published: (2025)
by: Rezk, Fady, et al.
Published: (2025)
AI Alignment with Changing and Influenceable Reward Functions
by: Carroll, Micah, et al.
Published: (2024)
by: Carroll, Micah, et al.
Published: (2024)
Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners
by: Muslimani, Calarina, et al.
Published: (2025)
by: Muslimani, Calarina, et al.
Published: (2025)
Finding Common Ground in a Sea of Alternatives
by: Chooi, Jay, et al.
Published: (2026)
by: Chooi, Jay, et al.
Published: (2026)
Pairwise Difference Learning for Classification
by: Belaid, Mohamed Karim, et al.
Published: (2024)
by: Belaid, Mohamed Karim, et al.
Published: (2024)
From Demonstrations to Rewards: Alignment Without Explicit Human Preferences
by: Zeng, Siliang, et al.
Published: (2025)
by: Zeng, Siliang, et al.
Published: (2025)
Expectation Alignment: Handling Reward Misspecification in the Presence of Expectation Mismatch
by: Mechergui, Malek, et al.
Published: (2024)
by: Mechergui, Malek, et al.
Published: (2024)
Similar Items
-
Axioms for AI Alignment from Human Feedback
by: Ge, Luise, et al.
Published: (2024) -
Strategic Classification With Externalities
by: Hossain, Safwan, et al.
Published: (2024) -
How RLHF Amplifies Sycophancy
by: Shapira, Itai, et al.
Published: (2026) -
Clone-Robust AI Alignment
by: Procaccia, Ariel D., et al.
Published: (2025) -
Generative Social Choice
by: Fish, Sara, et al.
Published: (2023)