Understanding Impact of Human Feedback via Influence Functions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Min, Taywon, Lee, Haeone, Kwon, Yongchan, Lee, Kimin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914014720688128
author Min, Taywon
Lee, Haeone
Kwon, Yongchan
Lee, Kimin
author_facet Min, Taywon
Lee, Haeone
Kwon, Yongchan
Lee, Kimin
contents In Reinforcement Learning from Human Feedback (RLHF), it is crucial to learn suitable reward models from human feedback to align large language models (LLMs) with human intentions. However, human feedback can often be noisy, inconsistent, or biased, especially when evaluating complex responses. Such feedback can lead to misaligned reward signals, potentially causing unintended side effects during the RLHF process. To address these challenges, we explore the use of influence functions to measure the impact of human feedback on the performance of reward models. We propose a compute-efficient approximation method that enables the application of influence functions to LLM-based reward models and large-scale preference datasets. Our experiments showcase two key applications of influence functions: (1) detecting common labeler biases in human feedback datasets and (2) guiding labelers in refining their strategies to better align with expert feedback. By quantifying the impact of human feedback, we believe that influence functions can enhance feedback interpretability and contribute to scalable oversight in RLHF, helping labelers provide more accurate and consistent feedback. Source code is available at https://github.com/mintaywon/IF_RLHF
format Preprint
id arxiv_https___arxiv_org_abs_2501_05790
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Understanding Impact of Human Feedback via Influence Functions
Min, Taywon
Lee, Haeone
Kwon, Yongchan
Lee, Kimin
Artificial Intelligence
Human-Computer Interaction
Machine Learning
In Reinforcement Learning from Human Feedback (RLHF), it is crucial to learn suitable reward models from human feedback to align large language models (LLMs) with human intentions. However, human feedback can often be noisy, inconsistent, or biased, especially when evaluating complex responses. Such feedback can lead to misaligned reward signals, potentially causing unintended side effects during the RLHF process. To address these challenges, we explore the use of influence functions to measure the impact of human feedback on the performance of reward models. We propose a compute-efficient approximation method that enables the application of influence functions to LLM-based reward models and large-scale preference datasets. Our experiments showcase two key applications of influence functions: (1) detecting common labeler biases in human feedback datasets and (2) guiding labelers in refining their strategies to better align with expert feedback. By quantifying the impact of human feedback, we believe that influence functions can enhance feedback interpretability and contribute to scalable oversight in RLHF, helping labelers provide more accurate and consistent feedback. Source code is available at https://github.com/mintaywon/IF_RLHF
title Understanding Impact of Human Feedback via Influence Functions
topic Artificial Intelligence
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2501.05790