A Common Pitfall of Margin-based Language Model Alignment: Gradient Entanglement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yuan, Hui, Zeng, Yifan, Wu, Yue, Wang, Huazheng, Wang, Mengdi, Leqi, Liu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
von: Qiu, Jiahao, et al.
Veröffentlicht: (2024)
von: Qiu, Jiahao, et al.
Veröffentlicht: (2024)
Self-Play with Adversarial Critic: Provable and Scalable Offline Alignment for Language Models
von: Ji, Xiang, et al.
Veröffentlicht: (2024)
von: Ji, Xiang, et al.
Veröffentlicht: (2024)
Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment
von: Zhang, Yifan, et al.
Veröffentlicht: (2024)
von: Zhang, Yifan, et al.
Veröffentlicht: (2024)
Larger or Smaller Reward Margins to Select Preferences for Alignment?
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
SAIL: Self-Improving Efficient Online Alignment of Large Language Models
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models
von: Bello, Femi, et al.
Veröffentlicht: (2025)
von: Bello, Femi, et al.
Veröffentlicht: (2025)
Self-Play Preference Optimization for Language Model Alignment
von: Wu, Yue, et al.
Veröffentlicht: (2024)
von: Wu, Yue, et al.
Veröffentlicht: (2024)
Personalized Language Modeling from Personalized Human Feedback
von: Li, Xinyu, et al.
Veröffentlicht: (2024)
von: Li, Xinyu, et al.
Veröffentlicht: (2024)
MaxMin-RLHF: Alignment with Diverse Human Preferences
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024)
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024)
Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective
von: Wang, Siwei, et al.
Veröffentlicht: (2025)
von: Wang, Siwei, et al.
Veröffentlicht: (2025)
Higher-order Linear Attention
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
AlignBench: Benchmarking Chinese Alignment of Large Language Models
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
Prompting Fairness: Integrating Causality to Debias Large Language Models
von: Li, Jingling, et al.
Veröffentlicht: (2024)
von: Li, Jingling, et al.
Veröffentlicht: (2024)
Probabilistic Token Alignment for Large Language Model Fusion
von: Zeng, Runjia, et al.
Veröffentlicht: (2025)
von: Zeng, Runjia, et al.
Veröffentlicht: (2025)
CARE-RFT: Confidence-Anchored Reinforcement Finetuning for Reliable Reasoning in Large Language Models
von: Li, Shuozhe, et al.
Veröffentlicht: (2026)
von: Li, Shuozhe, et al.
Veröffentlicht: (2026)
On the Role of Preference Variance in Preference Optimization
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
von: Abbes, Istabrak, et al.
Veröffentlicht: (2025)
von: Abbes, Istabrak, et al.
Veröffentlicht: (2025)
Training and Evaluating Language Models with Template-based Data Generation
von: Zhang, Yifan
Veröffentlicht: (2024)
von: Zhang, Yifan
Veröffentlicht: (2024)
CoRA: Optimizing Low-Rank Adaptation with Common Subspace of Large Language Models
von: Xiao, Xiaojun, et al.
Veröffentlicht: (2024)
von: Xiao, Xiaojun, et al.
Veröffentlicht: (2024)
Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models
von: Li, Chengao, et al.
Veröffentlicht: (2025)
von: Li, Chengao, et al.
Veröffentlicht: (2025)
SEE: Sememe Entanglement Encoding for Transformer-bases Models Compression
von: Zhang, Jing, et al.
Veröffentlicht: (2024)
von: Zhang, Jing, et al.
Veröffentlicht: (2024)
Accelerated Preference Optimization for Large Language Model Alignment
von: He, Jiafan, et al.
Veröffentlicht: (2024)
von: He, Jiafan, et al.
Veröffentlicht: (2024)
Curriculum Learning-Guided Progressive Distillation in Large Language Models
von: Cao, Jincheng, et al.
Veröffentlicht: (2026)
von: Cao, Jincheng, et al.
Veröffentlicht: (2026)
On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
FlashSampling: Fast and Memory-Efficient Exact Sampling
von: Ruiz, Tomas, et al.
Veröffentlicht: (2026)
von: Ruiz, Tomas, et al.
Veröffentlicht: (2026)
Propagation and Pitfalls: Reasoning-based Assessment of Knowledge Editing through Counterfactual Tasks
von: Hua, Wenyue, et al.
Veröffentlicht: (2024)
von: Hua, Wenyue, et al.
Veröffentlicht: (2024)
What Makes an Evaluation Useful? Common Pitfalls and Best Practices
von: Gekker, Gil, et al.
Veröffentlicht: (2025)
von: Gekker, Gil, et al.
Veröffentlicht: (2025)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
von: Wu, Junkang, et al.
Veröffentlicht: (2024)
Interactive Benchmarks
von: Yue, Baoqing, et al.
Veröffentlicht: (2026)
von: Yue, Baoqing, et al.
Veröffentlicht: (2026)
Temporal Consistency for LLM Reasoning Process Error Identification
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
von: Wei, Boyi, et al.
Veröffentlicht: (2024)
von: Wei, Boyi, et al.
Veröffentlicht: (2024)
Deep Delta Learning
von: Zhang, Yifan, et al.
Veröffentlicht: (2026)
von: Zhang, Yifan, et al.
Veröffentlicht: (2026)
SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths
von: Huang, Kaixuan, et al.
Veröffentlicht: (2024)
von: Huang, Kaixuan, et al.
Veröffentlicht: (2024)
Wanda++: Pruning Large Language Models via Regional Gradients
von: Yang, Yifan, et al.
Veröffentlicht: (2025)
von: Yang, Yifan, et al.
Veröffentlicht: (2025)
From Instructions to Constraints: Language Model Alignment with Automatic Constraint Verification
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs
von: Zeng, Yifan, et al.
Veröffentlicht: (2026)
von: Zeng, Yifan, et al.
Veröffentlicht: (2026)
Can Brain Signals Reveal Inner Alignment with Human Languages?
von: Han, William, et al.
Veröffentlicht: (2022)
von: Han, William, et al.
Veröffentlicht: (2022)
LLMs as Zero-shot Graph Learners: Alignment of GNN Representations with LLM Token Embeddings
von: Wang, Duo, et al.
Veröffentlicht: (2024)
von: Wang, Duo, et al.
Veröffentlicht: (2024)
A Markov Categorical Framework for Language Modeling
von: Zhang, Yifan
Veröffentlicht: (2025)
von: Zhang, Yifan
Veröffentlicht: (2025)
Ähnliche Einträge
-
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
von: Qiu, Jiahao, et al.
Veröffentlicht: (2024) -
Self-Play with Adversarial Critic: Provable and Scalable Offline Alignment for Language Models
von: Ji, Xiang, et al.
Veröffentlicht: (2024) -
Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment
von: Zhang, Yifan, et al.
Veröffentlicht: (2024) -
Larger or Smaller Reward Margins to Select Preferences for Alignment?
von: Huang, Kexin, et al.
Veröffentlicht: (2025) -
SAIL: Self-Improving Efficient Online Alignment of Large Language Models
von: Ding, Mucong, et al.
Veröffentlicht: (2024)