Plug-and-Play Training Framework for Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Jingyuan, Li, Rui, Li, Zheng, Sha, Lei, Sui, Zhifang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916545927577600
author Ma, Jingyuan
Li, Rui
Li, Zheng
Sha, Lei
Sui, Zhifang
author_facet Ma, Jingyuan
Li, Rui
Li, Zheng
Sha, Lei
Sui, Zhifang
contents Recently, preference optimization methods such as DPO have significantly enhanced large language models (LLMs) in wide tasks including dialogue and question-answering. However, current methods fail to account for the varying difficulty levels of training samples during preference optimization, leading to mediocre performance in tasks with high accuracy requirements, particularly in mathematical reasoning. To address this limitation, we propose a novel training framework, which employs multiple sampling to analyze output distributions, assign different weights to samples, and incorporate these weights into the preference optimization process. This plug-and-play approach enables LLMs to prioritize challenging examples during training, improving learning efficiency. Experimental results demonstrate that our framework integrates seamlessly with various preference optimization methods and achieves consistent improvements in mathematical reasoning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20996
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Plug-and-Play Training Framework for Preference Optimization
Ma, Jingyuan
Li, Rui
Li, Zheng
Sha, Lei
Sui, Zhifang
Computation and Language
Recently, preference optimization methods such as DPO have significantly enhanced large language models (LLMs) in wide tasks including dialogue and question-answering. However, current methods fail to account for the varying difficulty levels of training samples during preference optimization, leading to mediocre performance in tasks with high accuracy requirements, particularly in mathematical reasoning. To address this limitation, we propose a novel training framework, which employs multiple sampling to analyze output distributions, assign different weights to samples, and incorporate these weights into the preference optimization process. This plug-and-play approach enables LLMs to prioritize challenging examples during training, improving learning efficiency. Experimental results demonstrate that our framework integrates seamlessly with various preference optimization methods and achieves consistent improvements in mathematical reasoning tasks.
title Plug-and-Play Training Framework for Preference Optimization
topic Computation and Language
url https://arxiv.org/abs/2412.20996