FairPO: Robust Preference Optimization for Fair Multi-Label Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Mondal, Soumen Kumar, Chanda, Prateek, Varmora, Akshit, Ramakrishnan, Ganesh
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908678040322048
author Mondal, Soumen Kumar
Chanda, Prateek
Varmora, Akshit
Ramakrishnan, Ganesh
author_facet Mondal, Soumen Kumar
Chanda, Prateek
Varmora, Akshit
Ramakrishnan, Ganesh
contents Multi-label classification (MLC) often suffers from performance disparities across labels. We propose \textbf{FairPO}, a framework combining preference-based loss and group-robust optimization to improve fairness by targeting underperforming labels. FairPO partitions labels into a \textit{privileged} set for targeted improvement and a \textit{non-privileged} set to maintain baseline performance. For privileged labels, a DPO-inspired preference loss addresses hard examples by correcting ranking errors between true labels and their confusing counterparts. A constrained objective maintains performance for non-privileged labels, while a Group Robust Preference Optimization (GRPO) formulation adaptively balances both objectives to mitigate bias. We also demonstrate FairPO's versatility with reference-free variants using Contrastive (CPO) and Simple (SimPO) Preference Optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2505_02433
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FairPO: Robust Preference Optimization for Fair Multi-Label Learning
Mondal, Soumen Kumar
Chanda, Prateek
Varmora, Akshit
Ramakrishnan, Ganesh
Machine Learning
Artificial Intelligence
Multi-label classification (MLC) often suffers from performance disparities across labels. We propose \textbf{FairPO}, a framework combining preference-based loss and group-robust optimization to improve fairness by targeting underperforming labels. FairPO partitions labels into a \textit{privileged} set for targeted improvement and a \textit{non-privileged} set to maintain baseline performance. For privileged labels, a DPO-inspired preference loss addresses hard examples by correcting ranking errors between true labels and their confusing counterparts. A constrained objective maintains performance for non-privileged labels, while a Group Robust Preference Optimization (GRPO) formulation adaptively balances both objectives to mitigate bias. We also demonstrate FairPO's versatility with reference-free variants using Contrastive (CPO) and Simple (SimPO) Preference Optimization.
title FairPO: Robust Preference Optimization for Fair Multi-Label Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.02433