Uncovering Factor Level Preferences to Improve Human-Model Alignment

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Oh, Juhyun, Kim, Eunsu, Kim, Jiseon, Xu, Wenda, Cha, Inha, Wang, William Yang, Oh, Alice
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909904202104832
author Oh, Juhyun
Kim, Eunsu
Kim, Jiseon
Xu, Wenda
Cha, Inha
Wang, William Yang
Oh, Alice
author_facet Oh, Juhyun
Kim, Eunsu
Kim, Jiseon
Xu, Wenda
Cha, Inha
Wang, William Yang
Oh, Alice
contents Large language models (LLMs) often exhibit tendencies that diverge from human preferences, such as favoring certain writing styles or producing overly verbose outputs. While crucial for improvement, identifying the factors driving these misalignments remains challenging due to existing evaluation methods' reliance on coarse-grained comparisons and lack of explainability. To address this, we introduce PROFILE, an automated framework to uncover and measure factor-level preference alignment of humans and LLMs. Using PROFILE, we analyze preference alignment across three key tasks: summarization, instruction-following, and document-based QA. We find a significant discrepancy: while LLMs show poor factor-level alignment with human preferences when generating texts, they demonstrate strong alignment in discrimination tasks. We demonstrate how leveraging the identified generation-discrimination gap can be used to improve LLM alignment through multiple approaches, including fine-tuning with self-guidance. Our work highlights the value of factor-level analysis for identifying hidden misalignments and provides a practical framework for improving LLM-human preference alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2410_06965
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Uncovering Factor Level Preferences to Improve Human-Model Alignment
Oh, Juhyun
Kim, Eunsu
Kim, Jiseon
Xu, Wenda
Cha, Inha
Wang, William Yang
Oh, Alice
Computation and Language
Artificial Intelligence
Large language models (LLMs) often exhibit tendencies that diverge from human preferences, such as favoring certain writing styles or producing overly verbose outputs. While crucial for improvement, identifying the factors driving these misalignments remains challenging due to existing evaluation methods' reliance on coarse-grained comparisons and lack of explainability. To address this, we introduce PROFILE, an automated framework to uncover and measure factor-level preference alignment of humans and LLMs. Using PROFILE, we analyze preference alignment across three key tasks: summarization, instruction-following, and document-based QA. We find a significant discrepancy: while LLMs show poor factor-level alignment with human preferences when generating texts, they demonstrate strong alignment in discrimination tasks. We demonstrate how leveraging the identified generation-discrimination gap can be used to improve LLM alignment through multiple approaches, including fine-tuning with self-guidance. Our work highlights the value of factor-level analysis for identifying hidden misalignments and provides a practical framework for improving LLM-human preference alignment.
title Uncovering Factor Level Preferences to Improve Human-Model Alignment
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.06965