Aligning to What? Limits to RLHF Based Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Barnhart, Logan, Bafghi, Reza Akbarian, Becker, Stephen, Raissi, Maziar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Centerlines to Hemodynamics: Anisotropic RBF Decoders for Coronary Arteries
von: Bafghi, Reza Akbarian, et al.
Veröffentlicht: (2026)
von: Bafghi, Reza Akbarian, et al.
Veröffentlicht: (2026)
Test-Driven Agentic Framework for Reliable Robot Controller
von: Tripathi, Shivanshu, et al.
Veröffentlicht: (2026)
von: Tripathi, Shivanshu, et al.
Veröffentlicht: (2026)
MixDiff: Mixing Natural and Synthetic Images for Robust Self-Supervised Representations
von: Bafghi, Reza Akbarian, et al.
Veröffentlicht: (2024)
von: Bafghi, Reza Akbarian, et al.
Veröffentlicht: (2024)
Parameter Efficient Fine-tuning of Self-supervised ViTs without Catastrophic Forgetting
von: Bafghi, Reza Akbarian, et al.
Veröffentlicht: (2024)
von: Bafghi, Reza Akbarian, et al.
Veröffentlicht: (2024)
Fine Tuning without Catastrophic Forgetting via Selective Low Rank Adaptation
von: Bafghi, Reza Akbarian, et al.
Veröffentlicht: (2025)
von: Bafghi, Reza Akbarian, et al.
Veröffentlicht: (2025)
Where Did Your Model Learn That? Label-free Influence for Self-supervised Learning
von: Harilal, Nidhin, et al.
Veröffentlicht: (2024)
von: Harilal, Nidhin, et al.
Veröffentlicht: (2024)
Solving the Inverse Alignment Problem for Efficient RLHF
von: Krishna, Shambhavi, et al.
Veröffentlicht: (2024)
von: Krishna, Shambhavi, et al.
Veröffentlicht: (2024)
ChatGLM-RLHF: Practices of Aligning Large Language Models with Human Feedback
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)
von: Hou, Zhenyu, et al.
Veröffentlicht: (2024)
Why Is RLHF Alignment Shallow? A Gradient Analysis
von: Young, Robin
Veröffentlicht: (2026)
von: Young, Robin
Veröffentlicht: (2026)
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2025)
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2025)
More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
von: Li, Aaron J., et al.
Veröffentlicht: (2024)
von: Li, Aaron J., et al.
Veröffentlicht: (2024)
SenseAI: A Human-in-the-Loop Dataset for RLHF-Aligned Financial Sentiment Reasoning
von: Kabalisa, Berny
Veröffentlicht: (2026)
von: Kabalisa, Berny
Veröffentlicht: (2026)
MaxMin-RLHF: Alignment with Diverse Human Preferences
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024)
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024)
Understanding Tool-Augmented Agents for Lean Formalization: A Factorial Analysis
von: Zhang, Ke, et al.
Veröffentlicht: (2026)
von: Zhang, Ke, et al.
Veröffentlicht: (2026)
A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs
von: Srewa, Mahmoud, et al.
Veröffentlicht: (2025)
von: Srewa, Mahmoud, et al.
Veröffentlicht: (2025)
RLHF in an SFT Way: From Optimal Solution to Reward-Weighted Alignment
von: Du, Yuhao, et al.
Veröffentlicht: (2025)
von: Du, Yuhao, et al.
Veröffentlicht: (2025)
Proxy-RLHF: Decoupling Generation and Alignment in Large Language Model with Proxy
von: Zhu, Yu, et al.
Veröffentlicht: (2024)
von: Zhu, Yu, et al.
Veröffentlicht: (2024)
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
von: Ji, Jiaming, et al.
Veröffentlicht: (2024)
von: Ji, Jiaming, et al.
Veröffentlicht: (2024)
Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models
von: Dam, Harvey, et al.
Veröffentlicht: (2025)
von: Dam, Harvey, et al.
Veröffentlicht: (2025)
Culturally Adaptive Explainable LLM Assessment for Multilingual Information Disorder: A Human-in-the-Loop Approach
von: Jouneghani, Maziar Kianimoghadam
Veröffentlicht: (2026)
von: Jouneghani, Maziar Kianimoghadam
Veröffentlicht: (2026)
From RLHF to Direct Alignment: A Theoretical Unification of Preference Learning for Large Language Models
von: Raheja, Tarun, et al.
Veröffentlicht: (2026)
von: Raheja, Tarun, et al.
Veröffentlicht: (2026)
RLHF: A comprehensive Survey for Cultural, Multimodal and Low Latency Alignment Methods
von: Sharma, Raghav, et al.
Veröffentlicht: (2025)
von: Sharma, Raghav, et al.
Veröffentlicht: (2025)
RLHF Workflow: From Reward Modeling to Online RLHF
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models
von: Zheng, Chen, et al.
Veröffentlicht: (2025)
von: Zheng, Chen, et al.
Veröffentlicht: (2025)
Deep LPPLS: Forecasting of temporal critical points in natural, engineering and financial systems
von: Nielsen, Joshua, et al.
Veröffentlicht: (2024)
von: Nielsen, Joshua, et al.
Veröffentlicht: (2024)
Taming Overconfidence in LLMs: Reward Calibration in RLHF
von: Leng, Jixuan, et al.
Veröffentlicht: (2024)
von: Leng, Jixuan, et al.
Veröffentlicht: (2024)
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
von: Yu, Tianyu, et al.
Veröffentlicht: (2023)
von: Yu, Tianyu, et al.
Veröffentlicht: (2023)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
von: Hu, Jian, et al.
Veröffentlicht: (2024)
von: Hu, Jian, et al.
Veröffentlicht: (2024)
Language Models Learn to Mislead Humans via RLHF
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
TokAlign: Efficient Vocabulary Adaptation via Token Alignment
von: Li, Chong, et al.
Veröffentlicht: (2025)
von: Li, Chong, et al.
Veröffentlicht: (2025)
Divide-Then-Align: Honest Alignment based on the Knowledge Boundary of RAG
von: Sun, Xin, et al.
Veröffentlicht: (2025)
von: Sun, Xin, et al.
Veröffentlicht: (2025)
IterAlign: Iterative Constitutional Alignment of Large Language Models
von: Chen, Xiusi, et al.
Veröffentlicht: (2024)
von: Chen, Xiusi, et al.
Veröffentlicht: (2024)
SpeechAlign: a Framework for Speech Translation Alignment Evaluation
von: Alastruey, Belen, et al.
Veröffentlicht: (2023)
von: Alastruey, Belen, et al.
Veröffentlicht: (2023)
Failure Modes of Maximum Entropy RLHF
von: Çağatan, Ömer Veysel, et al.
Veröffentlicht: (2025)
von: Çağatan, Ömer Veysel, et al.
Veröffentlicht: (2025)
Physics-Informed Machine Learning for Smart Additive Manufacturing
von: Sharma, Rahul, et al.
Veröffentlicht: (2024)
von: Sharma, Rahul, et al.
Veröffentlicht: (2024)
Align-then-Unlearn: Embedding Alignment for LLM Unlearning
von: Spohn, Philipp, et al.
Veröffentlicht: (2025)
von: Spohn, Philipp, et al.
Veröffentlicht: (2025)
Fake Alignment: Are LLMs Really Aligned Well?
von: Wang, Yixu, et al.
Veröffentlicht: (2023)
von: Wang, Yixu, et al.
Veröffentlicht: (2023)
MKJ at SemEval-2026 Task 9: A Comparative Study of Generalist, Specialist, and Ensemble Strategies for Multilingual Polarization
von: Jouneghani, Maziar Kianimoghadam
Veröffentlicht: (2026)
von: Jouneghani, Maziar Kianimoghadam
Veröffentlicht: (2026)
CM-Align: Consistency-based Multilingual Alignment for Large Language Models
von: Zhang, Xue, et al.
Veröffentlicht: (2025)
von: Zhang, Xue, et al.
Veröffentlicht: (2025)
Align to the Pivot: Dual Alignment with Self-Feedback for Multilingual Math Reasoning
von: Zhao, Chunxu, et al.
Veröffentlicht: (2026)
von: Zhao, Chunxu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
From Centerlines to Hemodynamics: Anisotropic RBF Decoders for Coronary Arteries
von: Bafghi, Reza Akbarian, et al.
Veröffentlicht: (2026) -
Test-Driven Agentic Framework for Reliable Robot Controller
von: Tripathi, Shivanshu, et al.
Veröffentlicht: (2026) -
MixDiff: Mixing Natural and Synthetic Images for Robust Self-Supervised Representations
von: Bafghi, Reza Akbarian, et al.
Veröffentlicht: (2024) -
Parameter Efficient Fine-tuning of Self-supervised ViTs without Catastrophic Forgetting
von: Bafghi, Reza Akbarian, et al.
Veröffentlicht: (2024) -
Fine Tuning without Catastrophic Forgetting via Selective Low Rank Adaptation
von: Bafghi, Reza Akbarian, et al.
Veröffentlicht: (2025)