Aligning CodeLLMs with Direct Preference Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Miao, Yibo, Gao, Bofei, Quan, Shanghaoran, Lin, Junyang, Zan, Daoguang, Liu, Jiaheng, Yang, Jian, Liu, Tianyu, Deng, Zhijie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating and Aligning CodeLLMs on Human Preference
by: Yang, Jian, et al.
Published: (2024)
by: Yang, Jian, et al.
Published: (2024)
Mastering the Craft of Data Synthesis for CodeLLMs
by: Chen, Meng, et al.
Published: (2024)
by: Chen, Meng, et al.
Published: (2024)
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
by: Quan, Shanghaoran, et al.
Published: (2025)
by: Quan, Shanghaoran, et al.
Published: (2025)
Token-Importance Guided Direct Preference Optimization
by: Yang, Ning, et al.
Published: (2025)
by: Yang, Ning, et al.
Published: (2025)
Multi-Agent Collaboration for Multilingual Code Instruction Tuning
by: Yang, Jian, et al.
Published: (2025)
by: Yang, Jian, et al.
Published: (2025)
Automatically Generating Numerous Context-Driven SFT Data for LLMs across Diverse Granularity
by: Quan, Shanghaoran
Published: (2024)
by: Quan, Shanghaoran
Published: (2024)
Efficient Detection of LLM-generated Texts with a Bayesian Surrogate Model
by: Miao, Yibo, et al.
Published: (2023)
by: Miao, Yibo, et al.
Published: (2023)
Language Models can Self-Lengthen to Generate Long Texts
by: Quan, Shanghaoran, et al.
Published: (2024)
by: Quan, Shanghaoran, et al.
Published: (2024)
Towards a Unified View of Preference Learning for Large Language Models: A Survey
by: Gao, Bofei, et al.
Published: (2024)
by: Gao, Bofei, et al.
Published: (2024)
AdaMoE: Token-Adaptive Routing with Null Experts for Mixture-of-Experts Language Models
by: Zeng, Zihao, et al.
Published: (2024)
by: Zeng, Zihao, et al.
Published: (2024)
Learning the Boundary of Solvability: Aligning LLMs to Detect Unsolvable Problems
by: Peng, Dengyun, et al.
Published: (2025)
by: Peng, Dengyun, et al.
Published: (2025)
DMoERM: Recipes of Mixture-of-Experts for Effective Reward Modeling
by: Quan, Shanghaoran
Published: (2024)
by: Quan, Shanghaoran
Published: (2024)
Bayesian Exploration of Pre-trained Models for Low-shot Image Classification
by: Miao, Yibo, et al.
Published: (2024)
by: Miao, Yibo, et al.
Published: (2024)
Learning to Align Human Code Preferences
by: Yin, Xin, et al.
Published: (2025)
by: Yin, Xin, et al.
Published: (2025)
CodeV: Issue Resolving with Visual Data
by: Zhang, Linhao, et al.
Published: (2024)
by: Zhang, Linhao, et al.
Published: (2024)
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation
by: Yang, Gao, et al.
Published: (2025)
by: Yang, Gao, et al.
Published: (2025)
Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
by: Gao, Bofei, et al.
Published: (2024)
by: Gao, Bofei, et al.
Published: (2024)
DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs
by: Liu, Zhihan, et al.
Published: (2024)
by: Liu, Zhihan, et al.
Published: (2024)
Iterative Length-Regularized Direct Preference Optimization: A Case Study on Improving 7B Language Models to GPT-4 Level
by: Liu, Jie, et al.
Published: (2024)
by: Liu, Jie, et al.
Published: (2024)
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
by: Zhang, Shenao, et al.
Published: (2024)
by: Zhang, Shenao, et al.
Published: (2024)
Token-level Direct Preference Optimization
by: Zeng, Yongcheng, et al.
Published: (2024)
by: Zeng, Yongcheng, et al.
Published: (2024)
2D-DPO: Scaling Direct Preference Optimization with 2-Dimensional Supervision
by: Li, Shilong, et al.
Published: (2024)
by: Li, Shilong, et al.
Published: (2024)
InfiAlign: A Scalable and Sample-Efficient Framework for Aligning LLMs to Enhance Reasoning Capabilities
by: Cai, Shuo, et al.
Published: (2025)
by: Cai, Shuo, et al.
Published: (2025)
A GAN-based data poisoning framework against anomaly detection in vertical federated learning
by: Chen, Xiaolin, et al.
Published: (2024)
by: Chen, Xiaolin, et al.
Published: (2024)
COMAL: A Convergent Meta-Algorithm for Aligning LLMs with General Preferences
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
RPO-RAG: Aligning Small LLMs with Relation-aware Preference Optimization for Knowledge Graph Question Answering
by: Um, Kaehyun, et al.
Published: (2026)
by: Um, Kaehyun, et al.
Published: (2026)
Understanding How CodeLLMs (Mis)Predict Types with Activation Steering
by: Lucchetti, Francesca, et al.
Published: (2024)
by: Lucchetti, Francesca, et al.
Published: (2024)
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training
by: Wang, Pengkai, et al.
Published: (2025)
by: Wang, Pengkai, et al.
Published: (2025)
3D-Properties: Identifying Challenges in DPO and Charting a Path Forward
by: Yan, Yuzi, et al.
Published: (2024)
by: Yan, Yuzi, et al.
Published: (2024)
Optimizing LLMs with Direct Preferences: A Data Efficiency Perspective
by: Bernardelle, Pietro, et al.
Published: (2024)
by: Bernardelle, Pietro, et al.
Published: (2024)
Stable Preference Optimization: A Bilevel Approach to Catastrophic Preference Shift
by: Jian, Chengtao, et al.
Published: (2025)
by: Jian, Chengtao, et al.
Published: (2025)
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization
by: Nguyen, Thanh Thi, et al.
Published: (2025)
by: Nguyen, Thanh Thi, et al.
Published: (2025)
Vibe Checker: Aligning Code Evaluation with Human Preference
by: Zhong, Ming, et al.
Published: (2025)
by: Zhong, Ming, et al.
Published: (2025)
CodeS: Natural Language to Code Repository via Multi-Layer Sketch
by: Zan, Daoguang, et al.
Published: (2024)
by: Zan, Daoguang, et al.
Published: (2024)
FlipAttack: Jailbreak LLMs via Flipping
by: Liu, Yue, et al.
Published: (2024)
by: Liu, Yue, et al.
Published: (2024)
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
by: Zhou, Zhanhui, et al.
Published: (2023)
by: Zhou, Zhanhui, et al.
Published: (2023)
Orthogonal Finetuning for Direct Preference Optimization
by: Yang, Chenxu, et al.
Published: (2024)
by: Yang, Chenxu, et al.
Published: (2024)
SIPO: Stabilized and Improved Preference Optimization for Aligning Diffusion Models
by: Yang, Xiaomeng, et al.
Published: (2025)
by: Yang, Xiaomeng, et al.
Published: (2025)
Intelligently Weighting Multiple Reference Models for Direct Preference Optimization of LLMs
by: Wu, Skyler, et al.
Published: (2025)
by: Wu, Skyler, et al.
Published: (2025)
$β$-DPO: Direct Preference Optimization with Dynamic $β$
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
Similar Items
-
Evaluating and Aligning CodeLLMs on Human Preference
by: Yang, Jian, et al.
Published: (2024) -
Mastering the Craft of Data Synthesis for CodeLLMs
by: Chen, Meng, et al.
Published: (2024) -
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
by: Quan, Shanghaoran, et al.
Published: (2025) -
Token-Importance Guided Direct Preference Optimization
by: Yang, Ning, et al.
Published: (2025) -
Multi-Agent Collaboration for Multilingual Code Instruction Tuning
by: Yang, Jian, et al.
Published: (2025)