Expectation Alignment: Handling Reward Misspecification in the Presence of Expectation Mismatch
Fuente:
arXiv
Saved in:
| Main Authors: | Mechergui, Malek, Sreedharan, Sarath |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reinforcement Learning with Conditional Expectation Reward
by: Xiao, Changyi, et al.
Published: (2026)
by: Xiao, Changyi, et al.
Published: (2026)
FedSSG: Expectation-Gated and History-Aware Drift Alignment for Federated Learning
by: Zhou, Zhanting, et al.
Published: (2025)
by: Zhou, Zhanting, et al.
Published: (2025)
Neural Expectation Operators
by: Qi, Qian
Published: (2025)
by: Qi, Qian
Published: (2025)
A Unified Theory of $θ$-Expectations
by: Qi, Qian
Published: (2025)
by: Qi, Qian
Published: (2025)
Towards a Law of Iterated Expectations for Heuristic Estimators
by: Christiano, Paul, et al.
Published: (2024)
by: Christiano, Paul, et al.
Published: (2024)
A Mean-Field Theory of $Θ$-Expectations
by: Qi, Qian
Published: (2025)
by: Qi, Qian
Published: (2025)
An Expectation-Maximization Algorithm for Domain Adaptation in Gaussian Causal Models
by: Javidian, Mohammad Ali
Published: (2026)
by: Javidian, Mohammad Ali
Published: (2026)
Expectation Maximization Pseudo Labels
by: Xu, Moucheng, et al.
Published: (2023)
by: Xu, Moucheng, et al.
Published: (2023)
GEPO: Group Expectation Policy Optimization for Stable Heterogeneous Reinforcement Learning
by: Zhang, Han, et al.
Published: (2025)
by: Zhang, Han, et al.
Published: (2025)
Explaining Machine Learning Predictive Models through Conditional Expectation Methods
by: Ruiz-España, Silvia, et al.
Published: (2026)
by: Ruiz-España, Silvia, et al.
Published: (2026)
Global Sensitivity Analysis for Engineering Design Based on Individual Conditional Expectations
by: Palar, Pramudita Satria, et al.
Published: (2025)
by: Palar, Pramudita Satria, et al.
Published: (2025)
Safe RLHF Beyond Expectation: Stochastic Dominance for Universal Spectral Risk Control
by: Chittepu, Yaswanth, et al.
Published: (2026)
by: Chittepu, Yaswanth, et al.
Published: (2026)
Beyond Winning: Margin of Victory Relative to Expectation Unlocks Accurate Skill Ratings
by: Shorewala, Shivam, et al.
Published: (2025)
by: Shorewala, Shivam, et al.
Published: (2025)
Beyond Expectations: Learning with Stochastic Dominance Made Practical
by: Cen, Shicong, et al.
Published: (2024)
by: Cen, Shicong, et al.
Published: (2024)
Efficient Imitation under Misspecification
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
Bayesian Deep Learning Via Expectation Maximization and Turbo Deep Approximate Message Passing
by: Xu, Wei, et al.
Published: (2024)
by: Xu, Wei, et al.
Published: (2024)
Partial Identifiability and Misspecification in Inverse Reinforcement Learning
by: Skalse, Joar, et al.
Published: (2024)
by: Skalse, Joar, et al.
Published: (2024)
Bridge the Inference Gaps of Neural Processes via Expectation Maximization
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
Latent Variable Modeling in Multi-Agent Reinforcement Learning via Expectation-Maximization for UAV-Based Wildlife Protection
by: Taghavi, Mazyar, et al.
Published: (2025)
by: Taghavi, Mazyar, et al.
Published: (2025)
Vertical LoRA: Dense Expectation-Maximization Interpretation of Transformers
by: Fu, Zhuolin
Published: (2024)
by: Fu, Zhuolin
Published: (2024)
To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization
by: Wang, Haozhe, et al.
Published: (2025)
by: Wang, Haozhe, et al.
Published: (2025)
Priors Matter: Addressing Misspecification in Bayesian Deep Q-Learning
by: van der Vaart, Pascal R., et al.
Published: (2025)
by: van der Vaart, Pascal R., et al.
Published: (2025)
Pairwise Calibrated Rewards for Pluralistic Alignment
by: Halpern, Daniel, et al.
Published: (2025)
by: Halpern, Daniel, et al.
Published: (2025)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
by: Wang, Chaoqi, et al.
Published: (2025)
by: Wang, Chaoqi, et al.
Published: (2025)
LEC: Linear Expectation Constraints for Selection-Conditioned Risk Control in Selective Prediction and Routing Systems
by: Wang, Zhiyuan, et al.
Published: (2025)
by: Wang, Zhiyuan, et al.
Published: (2025)
Detecting Training Data of Large Language Models via Expectation Maximization
by: Kim, Gyuwan, et al.
Published: (2024)
by: Kim, Gyuwan, et al.
Published: (2024)
AI Alignment with Changing and Influenceable Reward Functions
by: Carroll, Micah, et al.
Published: (2024)
by: Carroll, Micah, et al.
Published: (2024)
The Reward Model Selection Crisis in Personalized Alignment
by: Rezk, Fady, et al.
Published: (2025)
by: Rezk, Fady, et al.
Published: (2025)
From Rights to Rites: Expectations Management in Smart-Home AI
by: Vishwarupe, Varad, et al.
Published: (2026)
by: Vishwarupe, Varad, et al.
Published: (2026)
Black Box Model Explanations and the Human Interpretability Expectations -- An Analysis in the Context of Homicide Prediction
by: Ribeiro, José, et al.
Published: (2022)
by: Ribeiro, José, et al.
Published: (2022)
Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners
by: Muslimani, Calarina, et al.
Published: (2025)
by: Muslimani, Calarina, et al.
Published: (2025)
Detecting Model Misspecification in Amortized Bayesian Inference with Neural Networks: An Extended Investigation
by: Schmitt, Marvin, et al.
Published: (2024)
by: Schmitt, Marvin, et al.
Published: (2024)
Is Elo Rating Reliable? A Study Under Model Misspecification
by: Tang, Shange, et al.
Published: (2025)
by: Tang, Shange, et al.
Published: (2025)
Epistemic Traps: Rational Misalignment Driven by Model Misspecification
by: Xu, Xingcheng, et al.
Published: (2026)
by: Xu, Xingcheng, et al.
Published: (2026)
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
by: Zhang, Shenao, et al.
Published: (2024)
by: Zhang, Shenao, et al.
Published: (2024)
From Demonstrations to Rewards: Alignment Without Explicit Human Preferences
by: Zeng, Siliang, et al.
Published: (2025)
by: Zeng, Siliang, et al.
Published: (2025)
Unrealized Expectations: Comparing AI Methods vs Classical Algorithms for Maximum Independent Set
by: Wu, Yikai, et al.
Published: (2025)
by: Wu, Yikai, et al.
Published: (2025)
DiffEM: Learning from Corrupted Data with Diffusion Models via Expectation Maximization
by: Hosseintabar, Danial, et al.
Published: (2025)
by: Hosseintabar, Danial, et al.
Published: (2025)
PALATE: Peculiar Application of the Law of Total Expectation to Enhance the Evaluation of Deep Generative Models
by: Dziarmaga, Tadeusz, et al.
Published: (2025)
by: Dziarmaga, Tadeusz, et al.
Published: (2025)
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning
by: Rajaram, Sara, et al.
Published: (2025)
by: Rajaram, Sara, et al.
Published: (2025)
Similar Items
-
Reinforcement Learning with Conditional Expectation Reward
by: Xiao, Changyi, et al.
Published: (2026) -
FedSSG: Expectation-Gated and History-Aware Drift Alignment for Federated Learning
by: Zhou, Zhanting, et al.
Published: (2025) -
Neural Expectation Operators
by: Qi, Qian
Published: (2025) -
A Unified Theory of $θ$-Expectations
by: Qi, Qian
Published: (2025) -
Towards a Law of Iterated Expectations for Heuristic Estimators
by: Christiano, Paul, et al.
Published: (2024)