Saved in:
| Main Authors: | Banerjee, Debangshu, Gopalan, Aditya |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2410.23726 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Reliable, Uncertainty-Aware Alignment
by: Banerjee, Debangshu, et al.
Published: (2025)
by: Banerjee, Debangshu, et al.
Published: (2025)
Bad Values but Good Behavior: Learning Highly Misspecified Bandits and MDPs
by: Banerjee, Debangshu, et al.
Published: (2023)
by: Banerjee, Debangshu, et al.
Published: (2023)
Why DPO is a Misspecified Estimator and How to Fix It
by: Gopalan, Aditya, et al.
Published: (2025)
by: Gopalan, Aditya, et al.
Published: (2025)
Data Shifts Hurt CoT: A Theoretical Study
by: Yin, Lang, et al.
Published: (2025)
by: Yin, Lang, et al.
Published: (2025)
Reliable Policy Iteration: Performance Robustness Across Architecture and Environment Perturbations
by: Eshwar, S. R., et al.
Published: (2025)
by: Eshwar, S. R., et al.
Published: (2025)
APPA: Adaptive Preference Pluralistic Alignment for Fair Federated RLHF of LLMs
by: Srewa, Mahmoud, et al.
Published: (2026)
by: Srewa, Mahmoud, et al.
Published: (2026)
Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic Framework
by: Kim, Kihyun, et al.
Published: (2025)
by: Kim, Kihyun, et al.
Published: (2025)
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
by: Kumarappan, Adarsh, et al.
Published: (2026)
by: Kumarappan, Adarsh, et al.
Published: (2026)
Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment
by: Yang, Zhiqin, et al.
Published: (2026)
by: Yang, Zhiqin, et al.
Published: (2026)
MaxMin-RLHF: Alignment with Diverse Human Preferences
by: Chakraborty, Souradip, et al.
Published: (2024)
by: Chakraborty, Souradip, et al.
Published: (2024)
Democratic Preference Alignment via Sortition-Weighted RLHF
by: Sana, Suvadip, et al.
Published: (2026)
by: Sana, Suvadip, et al.
Published: (2026)
A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO
by: Zhou, Xingyu, et al.
Published: (2025)
by: Zhou, Xingyu, et al.
Published: (2025)
Provably Mitigating Corruption, Overoptimization, and Verbosity Simultaneously in Offline and Online RLHF/DPO Alignment
by: Chen, Ziyi, et al.
Published: (2025)
by: Chen, Ziyi, et al.
Published: (2025)
Shifting Uncertainty to Critical Moments: Towards Reliable Uncertainty Quantification for VLA Model
by: Tang, Yanchuan, et al.
Published: (2026)
by: Tang, Yanchuan, et al.
Published: (2026)
Towards Reward Fairness in RLHF: From a Resource Allocation Perspective
by: Ouyang, Sheng, et al.
Published: (2025)
by: Ouyang, Sheng, et al.
Published: (2025)
Proxy-RLHF: Decoupling Generation and Alignment in Large Language Model with Proxy
by: Zhu, Yu, et al.
Published: (2024)
by: Zhu, Yu, et al.
Published: (2024)
RLHF in an SFT Way: From Optimal Solution to Reward-Weighted Alignment
by: Du, Yuhao, et al.
Published: (2025)
by: Du, Yuhao, et al.
Published: (2025)
Certifiable Safe RLHF: Fixed-Penalty Constraint Optimization for Safer Language Models
by: Pandit, Kartik, et al.
Published: (2025)
by: Pandit, Kartik, et al.
Published: (2025)
RLHF Workflow: From Reward Modeling to Online RLHF
by: Dong, Hanze, et al.
Published: (2024)
by: Dong, Hanze, et al.
Published: (2024)
ROCM: RLHF on consistency models
by: Shekhar, Shivanshu, et al.
Published: (2025)
by: Shekhar, Shivanshu, et al.
Published: (2025)
RLHF: A comprehensive Survey for Cultural, Multimodal and Low Latency Alignment Methods
by: Sharma, Raghav, et al.
Published: (2025)
by: Sharma, Raghav, et al.
Published: (2025)
HealthMamba: An Uncertainty-aware Spatiotemporal Graph State Space Model for Effective and Reliable Healthcare Facility Visit Prediction
by: Yu, Dahai, et al.
Published: (2026)
by: Yu, Dahai, et al.
Published: (2026)
Distributionally Robust Token Optimization in RLHF
by: Jin, Yeping, et al.
Published: (2026)
by: Jin, Yeping, et al.
Published: (2026)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
by: Hu, Jian, et al.
Published: (2024)
by: Hu, Jian, et al.
Published: (2024)
Towards Reliable Time Series Forecasting under Future Uncertainty: Ambiguity and Novelty Rejection Mechanisms
by: Feng, Ninghui, et al.
Published: (2025)
by: Feng, Ninghui, et al.
Published: (2025)
Towards Data-Centric RLHF: Simple Metrics for Preference Dataset Comparison
by: Shen, Judy Hanwen, et al.
Published: (2024)
by: Shen, Judy Hanwen, et al.
Published: (2024)
Boosting Deductive Reasoning with Step Signals In RLHF
by: Li, Jialian, et al.
Published: (2024)
by: Li, Jialian, et al.
Published: (2024)
The Hidden Link Between RLHF and Contrastive Learning
by: Lv, Xufei, et al.
Published: (2025)
by: Lv, Xufei, et al.
Published: (2025)
LLM Misalignment via Adversarial RLHF Platforms
by: Entezami, Erfan, et al.
Published: (2025)
by: Entezami, Erfan, et al.
Published: (2025)
CUAL: Continual Uncertainty-aware Active Learner
by: Rios, Amanda, et al.
Published: (2024)
by: Rios, Amanda, et al.
Published: (2024)
Conformal Feedback Alignment: Quantifying Answer-Level Reliability for Robust LLM Alignment
by: Chen, Tiejin, et al.
Published: (2026)
by: Chen, Tiejin, et al.
Published: (2026)
Reward-Robust RLHF in LLMs
by: Yan, Yuzi, et al.
Published: (2024)
by: Yan, Yuzi, et al.
Published: (2024)
RLHF and IIA: Perverse Incentives
by: Xu, Wanqiao, et al.
Published: (2023)
by: Xu, Wanqiao, et al.
Published: (2023)
Policy Filtration for RLHF to Mitigate Noise in Reward Models
by: Zhang, Chuheng, et al.
Published: (2024)
by: Zhang, Chuheng, et al.
Published: (2024)
A Descriptive and Normative Theory of Human Beliefs in RLHF
by: Dandekar, Sylee, et al.
Published: (2025)
by: Dandekar, Sylee, et al.
Published: (2025)
KL-regularization Itself is Differentially Private in Bandits and RLHF
by: Zhang, Yizhou, et al.
Published: (2025)
by: Zhang, Yizhou, et al.
Published: (2025)
Policy-labeled Preference Learning: Is Preference Enough for RLHF?
by: Cho, Taehyun, et al.
Published: (2025)
by: Cho, Taehyun, et al.
Published: (2025)
Uncertainty-aware Human Mobility Modeling and Anomaly Detection
by: Wen, Haomin, et al.
Published: (2024)
by: Wen, Haomin, et al.
Published: (2024)
Uncertainty-aware Traffic Prediction under Missing Data
by: Mei, Hao, et al.
Published: (2023)
by: Mei, Hao, et al.
Published: (2023)
Towards Label-Free Biological Reasoning Synthetic Dataset Creation via Uncertainty Filtering
by: Stoisser, Josefa Lia, et al.
Published: (2025)
by: Stoisser, Josefa Lia, et al.
Published: (2025)
Similar Items
-
Towards Reliable, Uncertainty-Aware Alignment
by: Banerjee, Debangshu, et al.
Published: (2025) -
Bad Values but Good Behavior: Learning Highly Misspecified Bandits and MDPs
by: Banerjee, Debangshu, et al.
Published: (2023) -
Why DPO is a Misspecified Estimator and How to Fix It
by: Gopalan, Aditya, et al.
Published: (2025) -
Data Shifts Hurt CoT: A Theoretical Study
by: Yin, Lang, et al.
Published: (2025) -
Reliable Policy Iteration: Performance Robustness Across Architecture and Environment Perturbations
by: Eshwar, S. R., et al.
Published: (2025)