Less is More: Improving LLM Alignment via Preference Data Selection
Fuente:
arXiv
Guardado en:
| Autores principales: | Deng, Xun, Zhong, Han, Ai, Rui, Feng, Fuli, Wang, Zheng, He, Xiangnan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Larger or Smaller Reward Margins to Select Preferences for Alignment?
por: Huang, Kexin, et al.
Publicado: (2025)
por: Huang, Kexin, et al.
Publicado: (2025)
Learn More, Forget Less: A Gradient-Aware Data Selection Approach for LLM
por: Liu, Yibai, et al.
Publicado: (2025)
por: Liu, Yibai, et al.
Publicado: (2025)
A3S: A General Active Clustering Method with Pairwise Constraints
por: Deng, Xun, et al.
Publicado: (2024)
por: Deng, Xun, et al.
Publicado: (2024)
Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
por: Zhang, Yuheng, et al.
Publicado: (2025)
por: Zhang, Yuheng, et al.
Publicado: (2025)
Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment
por: Yang, Rui, et al.
Publicado: (2024)
por: Yang, Rui, et al.
Publicado: (2024)
Less is More for Improving Automatic Evaluation of Factual Consistency
por: Wang, Tong, et al.
Publicado: (2024)
por: Wang, Tong, et al.
Publicado: (2024)
Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization
por: Wu, Junkang, et al.
Publicado: (2024)
por: Wu, Junkang, et al.
Publicado: (2024)
VISPA: Pluralistic Alignment via Automatic Value Selection and Activation
por: Zheng, Shenyan, et al.
Publicado: (2026)
por: Zheng, Shenyan, et al.
Publicado: (2026)
AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
por: Xiao, Jianfei, et al.
Publicado: (2026)
por: Xiao, Jianfei, et al.
Publicado: (2026)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
por: Kim, Dongyoung, et al.
Publicado: (2024)
por: Kim, Dongyoung, et al.
Publicado: (2024)
LIMR: Less is More for RL Scaling
por: Li, Xuefeng, et al.
Publicado: (2025)
por: Li, Xuefeng, et al.
Publicado: (2025)
Less is More: Denoising Knowledge Graphs For Retrieval Augmented Generation
por: Zheng, Yilun, et al.
Publicado: (2025)
por: Zheng, Yilun, et al.
Publicado: (2025)
Adversarial Preference Optimization: Enhancing Your Alignment via RM-LLM Game
por: Cheng, Pengyu, et al.
Publicado: (2023)
por: Cheng, Pengyu, et al.
Publicado: (2023)
Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards
por: Wang, Haoxiang, et al.
Publicado: (2024)
por: Wang, Haoxiang, et al.
Publicado: (2024)
Less is More: Local Intrinsic Dimensions of Contextual Language Models
por: Ruppik, Benjamin Matthias, et al.
Publicado: (2025)
por: Ruppik, Benjamin Matthias, et al.
Publicado: (2025)
Supernova: Achieving More with Less in Transformer Architectures
por: Tanase, Andrei-Valentin, et al.
Publicado: (2025)
por: Tanase, Andrei-Valentin, et al.
Publicado: (2025)
Less Noise, More Voice: Reinforcement Learning for Reasoning via Instruction Purification
por: Guo, Yiju, et al.
Publicado: (2026)
por: Guo, Yiju, et al.
Publicado: (2026)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
por: Yang, Jinming, et al.
Publicado: (2026)
por: Yang, Jinming, et al.
Publicado: (2026)
Less Peaky and More Accurate CTC Forced Alignment by Label Priors
por: Huang, Ruizhe, et al.
Publicado: (2024)
por: Huang, Ruizhe, et al.
Publicado: (2024)
Input-Time Scaling: Adding Noise and Irrelevance into Less-Is-More Drastically Improves Reasoning Performance and Efficiency
por: Huang, Rapheal, et al.
Publicado: (2025)
por: Huang, Rapheal, et al.
Publicado: (2025)
Less is More: Extreme Gradient Boost Rank-1 Adaption for Efficient Finetuning of LLMs
por: Zhang, Yifei, et al.
Publicado: (2024)
por: Zhang, Yifei, et al.
Publicado: (2024)
MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
por: Wei, Lai, et al.
Publicado: (2023)
por: Wei, Lai, et al.
Publicado: (2023)
Accelerated Preference Optimization for Large Language Model Alignment
por: He, Jiafan, et al.
Publicado: (2024)
por: He, Jiafan, et al.
Publicado: (2024)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
por: Wu, Junkang, et al.
Publicado: (2024)
por: Wu, Junkang, et al.
Publicado: (2024)
PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
por: Bobbili, Sarat Chandra, et al.
Publicado: (2025)
por: Bobbili, Sarat Chandra, et al.
Publicado: (2025)
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
por: Gao, Mingqi, et al.
Publicado: (2024)
por: Gao, Mingqi, et al.
Publicado: (2024)
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
por: Liu, Wei, et al.
Publicado: (2023)
por: Liu, Wei, et al.
Publicado: (2023)
Large Language Models are Learnable Planners for Long-Term Recommendation
por: Shi, Wentao, et al.
Publicado: (2024)
por: Shi, Wentao, et al.
Publicado: (2024)
Explain Less, Understand More: Jargon Detection via Personalized Parameter-Efficient Fine-tuning
por: Wu, Bohao, et al.
Publicado: (2025)
por: Wu, Bohao, et al.
Publicado: (2025)
When More is Less: Understanding Chain-of-Thought Length in LLMs
por: Wu, Yuyang, et al.
Publicado: (2025)
por: Wu, Yuyang, et al.
Publicado: (2025)
ComPO: Preference Alignment via Comparison Oracles
por: Chen, Peter, et al.
Publicado: (2025)
por: Chen, Peter, et al.
Publicado: (2025)
SEAL: Safety-enhanced Aligned LLM Fine-tuning via Bilevel Data Selection
por: Shen, Han, et al.
Publicado: (2024)
por: Shen, Han, et al.
Publicado: (2024)
DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment
por: Wedgwood, James, et al.
Publicado: (2026)
por: Wedgwood, James, et al.
Publicado: (2026)
Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation
por: Dong, Guanting, et al.
Publicado: (2024)
por: Dong, Guanting, et al.
Publicado: (2024)
Preference Leakage: A Contamination Problem in LLM-as-a-judge
por: Li, Dawei, et al.
Publicado: (2025)
por: Li, Dawei, et al.
Publicado: (2025)
Say Less, Mean More: Leveraging Pragmatics in Retrieval-Augmented Generation
por: Riaz, Haris, et al.
Publicado: (2025)
por: Riaz, Haris, et al.
Publicado: (2025)
Bridging the Gap Between Preference Alignment and Machine Unlearning
por: Feng, Xiaohua, et al.
Publicado: (2025)
por: Feng, Xiaohua, et al.
Publicado: (2025)
References Improve LLM Alignment in Non-Verifiable Domains
por: Shi, Kejian, et al.
Publicado: (2026)
por: Shi, Kejian, et al.
Publicado: (2026)
Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
por: Li, Yulin, et al.
Publicado: (2025)
por: Li, Yulin, et al.
Publicado: (2025)
ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment
por: Obbad, Elyas, et al.
Publicado: (2024)
por: Obbad, Elyas, et al.
Publicado: (2024)
Ejemplares similares
-
Larger or Smaller Reward Margins to Select Preferences for Alignment?
por: Huang, Kexin, et al.
Publicado: (2025) -
Learn More, Forget Less: A Gradient-Aware Data Selection Approach for LLM
por: Liu, Yibai, et al.
Publicado: (2025) -
A3S: A General Active Clustering Method with Pairwise Constraints
por: Deng, Xun, et al.
Publicado: (2024) -
Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
por: Zhang, Yuheng, et al.
Publicado: (2025) -
Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment
por: Yang, Rui, et al.
Publicado: (2024)