Towards a Unified View of Preference Learning for Large Language Models: A Survey
Fuente:
arXiv
Salvato in:
| Autori principali: | Gao, Bofei, Song, Feifan, Miao, Yibo, Cai, Zefan, Yang, Zhe, Chen, Liang, Hu, Helan, Xu, Runxin, Dong, Qingxiu, Zheng, Ce, Quan, Shanghaoran, Xiao, Wen, Zhang, Ge, Zan, Daoguang, Lu, Keming, Yu, Bowen, Liu, Dayiheng, Cui, Zeyu, Yang, Jian, Sha, Lei, Wang, Houfeng, Sui, Zhifang, Wang, Peiyi, Liu, Tianyu, Chang, Baobao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback
di: Gao, Bofei, et al.
Pubblicazione: (2024)
di: Gao, Bofei, et al.
Pubblicazione: (2024)
Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
di: Gao, Bofei, et al.
Pubblicazione: (2024)
di: Gao, Bofei, et al.
Pubblicazione: (2024)
Aligning CodeLLMs with Direct Preference Optimization
di: Miao, Yibo, et al.
Pubblicazione: (2024)
di: Miao, Yibo, et al.
Pubblicazione: (2024)
Decoding in Geometry: Alleviating Embedding-Space Crowding for Complex Reasoning
di: Yang, Yixin, et al.
Pubblicazione: (2026)
di: Yang, Yixin, et al.
Pubblicazione: (2026)
Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding
di: Xia, Heming, et al.
Pubblicazione: (2024)
di: Xia, Heming, et al.
Pubblicazione: (2024)
ICDPO: Effectively Borrowing Alignment Capability of Others via In-context Direct Preference Optimization
di: Song, Feifan, et al.
Pubblicazione: (2024)
di: Song, Feifan, et al.
Pubblicazione: (2024)
P-Aligner: Enabling Pre-Alignment of Language Models via Principled Instruction Synthesis
di: Song, Feifan, et al.
Pubblicazione: (2025)
di: Song, Feifan, et al.
Pubblicazione: (2025)
Language Models can Self-Lengthen to Generate Long Texts
di: Quan, Shanghaoran, et al.
Pubblicazione: (2024)
di: Quan, Shanghaoran, et al.
Pubblicazione: (2024)
Chain-of-Thought Tokens are Computer Program Variables
di: Zhu, Fangwei, et al.
Pubblicazione: (2025)
di: Zhu, Fangwei, et al.
Pubblicazione: (2025)
Reducing Hallucinations in Entity Abstract Summarization with Facts-Template Decomposition
di: Zhu, Fangwei, et al.
Pubblicazione: (2024)
di: Zhu, Fangwei, et al.
Pubblicazione: (2024)
Not All Demonstration Examples are Equally Beneficial: Reweighting Demonstration Examples for In-Context Learning
di: Yang, Zhe, et al.
Pubblicazione: (2023)
di: Yang, Zhe, et al.
Pubblicazione: (2023)
Utilizing Local Hierarchy with Adversarial Training for Hierarchical Text Classification
di: Wang, Zihan, et al.
Pubblicazione: (2024)
di: Wang, Zihan, et al.
Pubblicazione: (2024)
Improving the Robustness of Distantly-Supervised Named Entity Recognition via Uncertainty-Aware Teacher Learning and Student-Student Collaborative Learning
di: Si, Shuzheng, et al.
Pubblicazione: (2023)
di: Si, Shuzheng, et al.
Pubblicazione: (2023)
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
di: Quan, Shanghaoran, et al.
Pubblicazione: (2025)
di: Quan, Shanghaoran, et al.
Pubblicazione: (2025)
RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection
di: Yang, Yixin, et al.
Pubblicazione: (2025)
di: Yang, Yixin, et al.
Pubblicazione: (2025)
Can Large Multimodal Models Uncover Deep Semantics Behind Images?
di: Yang, Yixin, et al.
Pubblicazione: (2024)
di: Yang, Yixin, et al.
Pubblicazione: (2024)
Be a Multitude to Itself: A Prompt Evolution Framework for Red Teaming
di: Li, Rui, et al.
Pubblicazione: (2025)
di: Li, Rui, et al.
Pubblicazione: (2025)
PeriodicLoRA: Breaking the Low-Rank Bottleneck in LoRA Optimization
di: Meng, Xiangdi, et al.
Pubblicazione: (2024)
di: Meng, Xiangdi, et al.
Pubblicazione: (2024)
PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain
di: Chen, Liang, et al.
Pubblicazione: (2024)
di: Chen, Liang, et al.
Pubblicazione: (2024)
Mitigating Overthinking through Reasoning Shaping
di: Song, Feifan, et al.
Pubblicazione: (2025)
di: Song, Feifan, et al.
Pubblicazione: (2025)
A Survey on In-context Learning
di: Dong, Qingxiu, et al.
Pubblicazione: (2022)
di: Dong, Qingxiu, et al.
Pubblicazione: (2022)
Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding
di: Song, Feifan, et al.
Pubblicazione: (2025)
di: Song, Feifan, et al.
Pubblicazione: (2025)
How Far are LLMs from Being Our Digital Twins? A Benchmark for Persona-Based Behavior Chain Simulation
di: Li, Rui, et al.
Pubblicazione: (2025)
di: Li, Rui, et al.
Pubblicazione: (2025)
Towards Harmonized Uncertainty Estimation for Large Language Models
di: Li, Rui, et al.
Pubblicazione: (2025)
di: Li, Rui, et al.
Pubblicazione: (2025)
Self-Boosting Large Language Models with Synthetic Preference Data
di: Dong, Qingxiu, et al.
Pubblicazione: (2024)
di: Dong, Qingxiu, et al.
Pubblicazione: (2024)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
di: Cai, Zefan, et al.
Pubblicazione: (2024)
di: Cai, Zefan, et al.
Pubblicazione: (2024)
Qwen2.5-Coder Technical Report
di: Hui, Binyuan, et al.
Pubblicazione: (2024)
di: Hui, Binyuan, et al.
Pubblicazione: (2024)
Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
di: Yang, An, et al.
Pubblicazione: (2024)
di: Yang, An, et al.
Pubblicazione: (2024)
Rethinking Semantic Parsing for Large Language Models: Enhancing LLM Performance with Semantic Hints
di: An, Kaikai, et al.
Pubblicazione: (2024)
di: An, Kaikai, et al.
Pubblicazione: (2024)
SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
di: Li, Zheng, et al.
Pubblicazione: (2025)
di: Li, Zheng, et al.
Pubblicazione: (2025)
Preference Ranking Optimization for Human Alignment
di: Song, Feifan, et al.
Pubblicazione: (2023)
di: Song, Feifan, et al.
Pubblicazione: (2023)
Scaling Data Diversity for Fine-Tuning Language Models in Human Alignment
di: Song, Feifan, et al.
Pubblicazione: (2024)
di: Song, Feifan, et al.
Pubblicazione: (2024)
Multi-Agent Collaboration for Multilingual Code Instruction Tuning
di: Yang, Jian, et al.
Pubblicazione: (2025)
di: Yang, Jian, et al.
Pubblicazione: (2025)
Immobilized Microorganisms for Targeted Release in COD Reduction of Oilfield Wastewater and Heavy Oil Viscosity Reduction
di: Yuning Gong, et al.
Pubblicazione: (2025)
di: Yuning Gong, et al.
Pubblicazione: (2025)
Inference-Time Scaling for Generalist Reward Modeling
di: Liu, Zijun, et al.
Pubblicazione: (2025)
di: Liu, Zijun, et al.
Pubblicazione: (2025)
DMoERM: Recipes of Mixture-of-Experts for Effective Reward Modeling
di: Quan, Shanghaoran
Pubblicazione: (2024)
di: Quan, Shanghaoran
Pubblicazione: (2024)
Automatically Generating Numerous Context-Driven SFT Data for LLMs across Diverse Granularity
di: Quan, Shanghaoran
Pubblicazione: (2024)
di: Quan, Shanghaoran
Pubblicazione: (2024)
SWE-Mirror: Scaling Issue-Resolving Datasets by Mirroring Issues Across Repositories
di: Wang, Junhao, et al.
Pubblicazione: (2025)
di: Wang, Junhao, et al.
Pubblicazione: (2025)
Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
di: Li, Lei, et al.
Pubblicazione: (2024)
di: Li, Lei, et al.
Pubblicazione: (2024)
A GAN-based data poisoning framework against anomaly detection in vertical federated learning
di: Chen, Xiaolin, et al.
Pubblicazione: (2024)
di: Chen, Xiaolin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback
di: Gao, Bofei, et al.
Pubblicazione: (2024) -
Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
di: Gao, Bofei, et al.
Pubblicazione: (2024) -
Aligning CodeLLMs with Direct Preference Optimization
di: Miao, Yibo, et al.
Pubblicazione: (2024) -
Decoding in Geometry: Alleviating Embedding-Space Crowding for Complex Reasoning
di: Yang, Yixin, et al.
Pubblicazione: (2026) -
Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding
di: Xia, Heming, et al.
Pubblicazione: (2024)