Understanding Likelihood Over-optimisation in Direct Alignment Algorithms
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Zhengyan, Land, Sander, Locatelli, Acyr, Geist, Matthieu, Bartolo, Max |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Benchmarking Is Just Feature Selection and Multiple Regression
by: Bowyer, Sam, et al.
Published: (2026)
by: Bowyer, Sam, et al.
Published: (2026)
Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts
by: Gritsch, Nikolas, et al.
Published: (2024)
by: Gritsch, Nikolas, et al.
Published: (2024)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
by: Rafailov, Rafael, et al.
Published: (2024)
by: Rafailov, Rafael, et al.
Published: (2024)
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
by: Razin, Noam, et al.
Published: (2024)
by: Razin, Noam, et al.
Published: (2024)
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
by: Agarwal, Rishabh, et al.
Published: (2023)
by: Agarwal, Rishabh, et al.
Published: (2023)
Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models
by: Land, Sander, et al.
Published: (2024)
by: Land, Sander, et al.
Published: (2024)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
by: Kim, Dongyoung, et al.
Published: (2024)
by: Kim, Dongyoung, et al.
Published: (2024)
Teaching Your Models to Understand Code via Focal Preference Alignment
by: Wu, Jie, et al.
Published: (2025)
by: Wu, Jie, et al.
Published: (2025)
Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
by: Goel, Raghavv, et al.
Published: (2024)
by: Goel, Raghavv, et al.
Published: (2024)
Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
by: Ruis, Laura, et al.
Published: (2024)
by: Ruis, Laura, et al.
Published: (2024)
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
by: Chen, Jianhui, et al.
Published: (2024)
by: Chen, Jianhui, et al.
Published: (2024)
Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation
by: Dong, Guanting, et al.
Published: (2024)
by: Dong, Guanting, et al.
Published: (2024)
Log-Likelihood, Simpson's Paradox, and the Detection of Machine-Generated Text
by: Kempton, Tom, et al.
Published: (2026)
by: Kempton, Tom, et al.
Published: (2026)
Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards
by: Wang, Haoxiang, et al.
Published: (2024)
by: Wang, Haoxiang, et al.
Published: (2024)
GDS Agent for Graph Algorithmic Reasoning
by: Shi, Borun, et al.
Published: (2025)
by: Shi, Borun, et al.
Published: (2025)
Would I Lie To You? Inference Time Alignment of Language Models using Direct Preference Heads
by: Hadji-Kyriacou, Avelina Asada, et al.
Published: (2024)
by: Hadji-Kyriacou, Avelina Asada, et al.
Published: (2024)
Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
by: Huang, Audrey, et al.
Published: (2024)
by: Huang, Audrey, et al.
Published: (2024)
Resolving UnderEdit & OverEdit with Iterative & Neighbor-Assisted Model Editing
by: Baghel, Bhiman Kumar, et al.
Published: (2025)
by: Baghel, Bhiman Kumar, et al.
Published: (2025)
Bridging the Creativity Understanding Gap: Small-Scale Human Alignment Enables Expert-Level Humor Ranking in LLMs
by: Zhou, Kuan Lok, et al.
Published: (2025)
by: Zhou, Kuan Lok, et al.
Published: (2025)
References Improve LLM Alignment in Non-Verifiable Domains
by: Shi, Kejian, et al.
Published: (2026)
by: Shi, Kejian, et al.
Published: (2026)
Tree-OPO: Off-policy Monte Carlo Tree-Guided Advantage Optimization for Multistep Reasoning
by: Huang, Bingning, et al.
Published: (2025)
by: Huang, Bingning, et al.
Published: (2025)
Reformatted Alignment
by: Fan, Run-Ze, et al.
Published: (2024)
by: Fan, Run-Ze, et al.
Published: (2024)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
Direct Behavior Optimization: Unlocking the Potential of Lightweight LLMs
by: Yang, Hongming, et al.
Published: (2025)
by: Yang, Hongming, et al.
Published: (2025)
Imitating Language via Scalable Inverse Reinforcement Learning
by: Wulfmeier, Markus, et al.
Published: (2024)
by: Wulfmeier, Markus, et al.
Published: (2024)
Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
by: Patel, Dev, et al.
Published: (2025)
by: Patel, Dev, et al.
Published: (2025)
Optimising Language Models for Downstream Tasks: A Post-Training Perspective
by: Shi, Zhengyan
Published: (2025)
by: Shi, Zhengyan
Published: (2025)
Tracing Facts or just Copies? A critical investigation of the Competitions of Mechanisms in Large Language Models
by: Campregher, Dante, et al.
Published: (2025)
by: Campregher, Dante, et al.
Published: (2025)
Model Extrapolation Expedites Alignment
by: Zheng, Chujie, et al.
Published: (2024)
by: Zheng, Chujie, et al.
Published: (2024)
Revisiting the Superficial Alignment Hypothesis
by: Raghavendra, Mohit, et al.
Published: (2024)
by: Raghavendra, Mohit, et al.
Published: (2024)
Sample-Efficient Alignment for LLMs
by: Liu, Zichen, et al.
Published: (2024)
by: Liu, Zichen, et al.
Published: (2024)
Variational Best-of-N Alignment
by: Amini, Afra, et al.
Published: (2024)
by: Amini, Afra, et al.
Published: (2024)
Aligners: Decoupling LLMs and Alignment
by: Ngweta, Lilian, et al.
Published: (2024)
by: Ngweta, Lilian, et al.
Published: (2024)
Continual Learning with Global Alignment
by: Bai, Xueying, et al.
Published: (2022)
by: Bai, Xueying, et al.
Published: (2022)
Test-Time Safety Alignment
by: Saglam, Baturay, et al.
Published: (2026)
by: Saglam, Baturay, et al.
Published: (2026)
Context Over Content: Exposing Evaluation Faking in Automated Judges
by: Gupta, Manan, et al.
Published: (2026)
by: Gupta, Manan, et al.
Published: (2026)
Agent-Omni: Test-Time Multimodal Reasoning via Model Coordination for Understanding Anything
by: Lin, Huawei, et al.
Published: (2025)
by: Lin, Huawei, et al.
Published: (2025)
Revisiting Graph-Tokenizing Large Language Models: A Systematic Evaluation of Graph Token Understanding
by: Zhang, Zhongjian, et al.
Published: (2026)
by: Zhang, Zhongjian, et al.
Published: (2026)
Alignment faking in large language models
by: Greenblatt, Ryan, et al.
Published: (2024)
by: Greenblatt, Ryan, et al.
Published: (2024)
Similar Items
-
Efficient Benchmarking Is Just Feature Selection and Multiple Regression
by: Bowyer, Sam, et al.
Published: (2026) -
Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts
by: Gritsch, Nikolas, et al.
Published: (2024) -
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
by: Rafailov, Rafael, et al.
Published: (2024) -
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
by: Razin, Noam, et al.
Published: (2024) -
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
by: Agarwal, Rishabh, et al.
Published: (2023)