One Goal, Many Challenges: Robust Preference Optimization Amid Content-Aware and Multi-Source Noise

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Afzali, Amirabbas, Afsharrad, Amirhossein, Mousavi, Seyed Shahabeddin, Lall, Sanjay
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912586754162688
author Afzali, Amirabbas
Afsharrad, Amirhossein
Mousavi, Seyed Shahabeddin
Lall, Sanjay
author_facet Afzali, Amirabbas
Afsharrad, Amirhossein
Mousavi, Seyed Shahabeddin
Lall, Sanjay
contents Large Language Models (LLMs) have made significant strides in generating human-like responses, largely due to preference alignment techniques. However, these methods often assume unbiased human feedback, which is rarely the case in real-world scenarios. This paper introduces Content-Aware Noise-Resilient Preference Optimization (CNRPO), a novel framework that addresses multiple sources of content-dependent noise in preference learning. CNRPO employs a multi-objective optimization approach to separate true preferences from content-aware noises, effectively mitigating their impact. We leverage backdoor attack mechanisms to efficiently learn and control various noise sources within a single model. Theoretical analysis and extensive experiments on different synthetic noisy datasets demonstrate that CNRPO significantly improves alignment with primary human preferences while controlling for secondary noises and biases, such as response length and harmfulness.
format Preprint
id arxiv_https___arxiv_org_abs_2503_12301
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle One Goal, Many Challenges: Robust Preference Optimization Amid Content-Aware and Multi-Source Noise
Afzali, Amirabbas
Afsharrad, Amirhossein
Mousavi, Seyed Shahabeddin
Lall, Sanjay
Machine Learning
Computation and Language
Large Language Models (LLMs) have made significant strides in generating human-like responses, largely due to preference alignment techniques. However, these methods often assume unbiased human feedback, which is rarely the case in real-world scenarios. This paper introduces Content-Aware Noise-Resilient Preference Optimization (CNRPO), a novel framework that addresses multiple sources of content-dependent noise in preference learning. CNRPO employs a multi-objective optimization approach to separate true preferences from content-aware noises, effectively mitigating their impact. We leverage backdoor attack mechanisms to efficiently learn and control various noise sources within a single model. Theoretical analysis and extensive experiments on different synthetic noisy datasets demonstrate that CNRPO significantly improves alignment with primary human preferences while controlling for secondary noises and biases, such as response length and harmfulness.
title One Goal, Many Challenges: Robust Preference Optimization Amid Content-Aware and Multi-Source Noise
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2503.12301