InfoPO: On Mutual Information Maximization for Large Language Model Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Teng, Ge, Zhen, Sanghavi, Sujay, Wang, Tian, Katz-Samuels, Julian, Versage, Marc, Cui, Qingjun, Chilimbi, Trishul
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916735135776768
author Xiao, Teng
Ge, Zhen
Sanghavi, Sujay
Wang, Tian
Katz-Samuels, Julian
Versage, Marc
Cui, Qingjun
Chilimbi, Trishul
author_facet Xiao, Teng
Ge, Zhen
Sanghavi, Sujay
Wang, Tian
Katz-Samuels, Julian
Versage, Marc
Cui, Qingjun
Chilimbi, Trishul
contents We study the post-training of large language models (LLMs) with human preference data. Recently, direct preference optimization and its variants have shown considerable promise in aligning language models, eliminating the need for reward models and online sampling. Despite these benefits, these methods rely on explicit assumptions about the Bradley-Terry (BT) model, which makes them prone to overfitting and results in suboptimal performance, particularly on reasoning-heavy tasks. To address these challenges, we propose a principled preference fine-tuning algorithm called InfoPO, which effectively and efficiently aligns large language models using preference data. InfoPO eliminates the reliance on the BT model and prevents the likelihood of the chosen response from decreasing. Extensive experiments confirm that InfoPO consistently outperforms established baselines on widely used open benchmarks, particularly in reasoning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_08507
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InfoPO: On Mutual Information Maximization for Large Language Model Alignment
Xiao, Teng
Ge, Zhen
Sanghavi, Sujay
Wang, Tian
Katz-Samuels, Julian
Versage, Marc
Cui, Qingjun
Chilimbi, Trishul
Machine Learning
We study the post-training of large language models (LLMs) with human preference data. Recently, direct preference optimization and its variants have shown considerable promise in aligning language models, eliminating the need for reward models and online sampling. Despite these benefits, these methods rely on explicit assumptions about the Bradley-Terry (BT) model, which makes them prone to overfitting and results in suboptimal performance, particularly on reasoning-heavy tasks. To address these challenges, we propose a principled preference fine-tuning algorithm called InfoPO, which effectively and efficiently aligns large language models using preference data. InfoPO eliminates the reliance on the BT model and prevents the likelihood of the chosen response from decreasing. Extensive experiments confirm that InfoPO consistently outperforms established baselines on widely used open benchmarks, particularly in reasoning tasks.
title InfoPO: On Mutual Information Maximization for Large Language Model Alignment
topic Machine Learning
url https://arxiv.org/abs/2505.08507