Leveraging Robust Optimization for LLM Alignment under Distribution Shifts
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Mingye, Liu, Yi, Fu, Zheren, Zhang, Yongdong, Mao, Zhendong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
In-Token Rationality Optimization: Towards Accurate and Concise LLM Reasoning via Self-Feedback
by: Zhu, Mingye, et al.
Published: (2025)
by: Zhu, Mingye, et al.
Published: (2025)
Leveraging Importance Sampling to Detach Alignment Modules from Large Language Models
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
DACL-RAG: Data Augmentation Strategy with Curriculum Learning for Retrieval-Augmented Generation
by: Wang, Shaohan, et al.
Published: (2025)
by: Wang, Shaohan, et al.
Published: (2025)
FlipGuard: Defending Preference Alignment against Update Regression with Constrained Optimization
by: Zhu, Mingye, et al.
Published: (2024)
by: Zhu, Mingye, et al.
Published: (2024)
On-the-fly Preference Alignment via Principle-Guided Decoding
by: Zhu, Mingye, et al.
Published: (2025)
by: Zhu, Mingye, et al.
Published: (2025)
SparseRM: A Lightweight Preference Modeling with Sparse Autoencoder
by: Liu, Dengcan, et al.
Published: (2025)
by: Liu, Dengcan, et al.
Published: (2025)
LIRE: listwise reward enhancement for preference alignment
by: Zhu, Mingye, et al.
Published: (2024)
by: Zhu, Mingye, et al.
Published: (2024)
Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models
by: Zhang, Huatian, et al.
Published: (2026)
by: Zhang, Huatian, et al.
Published: (2026)
Training LLM-Based Agents with Synthetic Self-Reflected Trajectories and Partial Masking
by: Chen, Yihan, et al.
Published: (2025)
by: Chen, Yihan, et al.
Published: (2025)
Align Documents to Questions: Question-Oriented Document Rewriting for Retrieval-Augmented Generation
by: Li, Jiaang, et al.
Published: (2026)
by: Li, Jiaang, et al.
Published: (2026)
Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models
by: Xia, Hou, et al.
Published: (2025)
by: Xia, Hou, et al.
Published: (2025)
FS-Researcher: Test-Time Scaling for Long-Horizon Research Tasks with File-System-Based Agents
by: Zhu, Chiwei, et al.
Published: (2026)
by: Zhu, Chiwei, et al.
Published: (2026)
ELDER: Enhancing Lifelong Model Editing with Mixture-of-LoRA
by: Li, Jiaang, et al.
Published: (2024)
by: Li, Jiaang, et al.
Published: (2024)
Wiki Live Challenge: Challenging Deep Research Agents with Expert-Level Wikipedia Articles
by: Wang, Shaohan, et al.
Published: (2026)
by: Wang, Shaohan, et al.
Published: (2026)
Benchmarking Large Language Models on Controllable Generation under Diversified Instructions
by: Chen, Yihan, et al.
Published: (2024)
by: Chen, Yihan, et al.
Published: (2024)
ExpertPrompting: Instructing Large Language Models to be Distinguished Experts
by: Xu, Benfeng, et al.
Published: (2023)
by: Xu, Benfeng, et al.
Published: (2023)
DPO-Shift: Shifting the Distribution of Direct Preference Optimization
by: Yang, Xiliang, et al.
Published: (2025)
by: Yang, Xiliang, et al.
Published: (2025)
Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts
by: Yin, Yueqin, et al.
Published: (2024)
by: Yin, Yueqin, et al.
Published: (2024)
Mitigating Biases in Language Models via Bias Unlearning
by: Liu, Dianqing, et al.
Published: (2025)
by: Liu, Dianqing, et al.
Published: (2025)
Robust Prompt Optimization for Large Language Models Against Distribution Shifts
by: Li, Moxin, et al.
Published: (2023)
by: Li, Moxin, et al.
Published: (2023)
Acceleration Multiple Heads Decoding for LLM via Dynamic Tree Attention
by: Zhang, Zhendong
Published: (2025)
by: Zhang, Zhendong
Published: (2025)
MetaRM: Shifted Distributions Alignment via Meta-Learning
by: Dou, Shihan, et al.
Published: (2024)
by: Dou, Shihan, et al.
Published: (2024)
ChiMed-GPT: A Chinese Medical Large Language Model with Full Training Regime and Better Alignment to Human Preferences
by: Tian, Yuanhe, et al.
Published: (2023)
by: Tian, Yuanhe, et al.
Published: (2023)
Enhancing Persona Following at Decoding Time via Dynamic Importance Estimation for Role-Playing Agents
by: Liu, Yuxin, et al.
Published: (2026)
by: Liu, Yuxin, et al.
Published: (2026)
Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization
by: Duan, Shitong, et al.
Published: (2024)
by: Duan, Shitong, et al.
Published: (2024)
From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding
by: Zhu, Chiwei, et al.
Published: (2025)
by: Zhu, Chiwei, et al.
Published: (2025)
Towards Robust Multimodal Emotion Recognition under Missing Modalities and Distribution Shifts
by: Zhong, Guowei, et al.
Published: (2025)
by: Zhong, Guowei, et al.
Published: (2025)
Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
by: Xu, Shusheng, et al.
Published: (2024)
by: Xu, Shusheng, et al.
Published: (2024)
Self-Augmented Preference Optimization: Off-Policy Paradigms for Language Model Alignment
by: Yin, Yueqin, et al.
Published: (2024)
by: Yin, Yueqin, et al.
Published: (2024)
Robust Preference Alignment via Directional Neighborhood Consensus
by: Mao, Ruochen, et al.
Published: (2025)
by: Mao, Ruochen, et al.
Published: (2025)
RCEM: Embedder Equipped with Query Rewriting Skill for Robust Conversational Search in Distributional Shift
by: Son, Kilho, et al.
Published: (2026)
by: Son, Kilho, et al.
Published: (2026)
Selection of LLM Fine-Tuning Data based on Orthogonal Rules
by: Li, Xiaomin, et al.
Published: (2024)
by: Li, Xiaomin, et al.
Published: (2024)
StoicLLM: Preference Optimization for Philosophical Alignment in Small Language Models
by: Khan, Ishmam, et al.
Published: (2026)
by: Khan, Ishmam, et al.
Published: (2026)
SPA: Achieving Consensus in LLM Alignment via Self-Priority Optimization
by: Huang, Yue, et al.
Published: (2025)
by: Huang, Yue, et al.
Published: (2025)
Measuring Distribution Shift in User Prompts and Its Effects on LLM Performance
by: Seegmiller, Parker, et al.
Published: (2026)
by: Seegmiller, Parker, et al.
Published: (2026)
Optimizing Alignment with Less: Leveraging Data Augmentation for Personalized Evaluation
by: Seraj, Javad, et al.
Published: (2024)
by: Seraj, Javad, et al.
Published: (2024)
InCo-DPO: Balancing Distribution Shift and Data Quality for Enhanced Preference Optimization
by: Wang, Yunan, et al.
Published: (2025)
by: Wang, Yunan, et al.
Published: (2025)
GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents
by: Li, Xueyi, et al.
Published: (2026)
by: Li, Xueyi, et al.
Published: (2026)
Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization
by: Zhou, Hongli, et al.
Published: (2026)
by: Zhou, Hongli, et al.
Published: (2026)
Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
Similar Items
-
In-Token Rationality Optimization: Towards Accurate and Concise LLM Reasoning via Self-Feedback
by: Zhu, Mingye, et al.
Published: (2025) -
Leveraging Importance Sampling to Detach Alignment Modules from Large Language Models
by: Liu, Yi, et al.
Published: (2025) -
DACL-RAG: Data Augmentation Strategy with Curriculum Learning for Retrieval-Augmented Generation
by: Wang, Shaohan, et al.
Published: (2025) -
FlipGuard: Defending Preference Alignment against Update Regression with Constrained Optimization
by: Zhu, Mingye, et al.
Published: (2024) -
On-the-fly Preference Alignment via Principle-Guided Decoding
by: Zhu, Mingye, et al.
Published: (2025)