WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Taiwei, Wang, Zhuoer, Yang, Longqi, Lin, Ying-Chun, He, Zexue, Wan, Mengting, Zhou, Pei, Jauhar, Sujay, Chen, Sihao, Xia, Shan, Zhang, Hongfei, Zhao, Jieyu, Xu, Xiaofeng, Song, Xia, Neville, Jennifer
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914483648069632
author Shi, Taiwei
Wang, Zhuoer
Yang, Longqi
Lin, Ying-Chun
He, Zexue
Wan, Mengting
Zhou, Pei
Jauhar, Sujay
Chen, Sihao
Xia, Shan
Zhang, Hongfei
Zhao, Jieyu
Xu, Xiaofeng
Song, Xia
Neville, Jennifer
author_facet Shi, Taiwei
Wang, Zhuoer
Yang, Longqi
Lin, Ying-Chun
He, Zexue
Wan, Mengting
Zhou, Pei
Jauhar, Sujay
Chen, Sihao
Xia, Shan
Zhang, Hongfei
Zhao, Jieyu
Xu, Xiaofeng
Song, Xia
Neville, Jennifer
contents As large language models (LLMs) continue to advance, aligning these models with human preferences has emerged as a critical challenge. Traditional alignment methods, relying on human or LLM annotated datasets, are limited by their resource-intensive nature, inherent subjectivity, misalignment with real-world user preferences, and the risk of feedback loops that amplify model biases. To overcome these limitations, we introduce WildFeedback, a novel framework that leverages in-situ user feedback during conversations with LLMs to create preference datasets automatically. Given a corpus of multi-turn user-LLM conversation, WildFeedback identifies and classifies user feedback to LLM responses between conversation turns. The user feedback is then used to create examples of preferred and dispreferred responses according to users' preference. Our experiments demonstrate that LLMs fine-tuned on WildFeedback dataset exhibit significantly improved alignment with user preferences, as evidenced by both traditional benchmarks and our proposed checklist-guided evaluation. By incorporating in-situ feedback from actual users, WildFeedback addresses the scalability, subjectivity, and bias challenges that plague existing approaches, marking a significant step toward developing LLMs that are more responsive to the diverse and evolving needs of their users.
format Preprint
id arxiv_https___arxiv_org_abs_2408_15549
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback
Shi, Taiwei
Wang, Zhuoer
Yang, Longqi
Lin, Ying-Chun
He, Zexue
Wan, Mengting
Zhou, Pei
Jauhar, Sujay
Chen, Sihao
Xia, Shan
Zhang, Hongfei
Zhao, Jieyu
Xu, Xiaofeng
Song, Xia
Neville, Jennifer
Computation and Language
As large language models (LLMs) continue to advance, aligning these models with human preferences has emerged as a critical challenge. Traditional alignment methods, relying on human or LLM annotated datasets, are limited by their resource-intensive nature, inherent subjectivity, misalignment with real-world user preferences, and the risk of feedback loops that amplify model biases. To overcome these limitations, we introduce WildFeedback, a novel framework that leverages in-situ user feedback during conversations with LLMs to create preference datasets automatically. Given a corpus of multi-turn user-LLM conversation, WildFeedback identifies and classifies user feedback to LLM responses between conversation turns. The user feedback is then used to create examples of preferred and dispreferred responses according to users' preference. Our experiments demonstrate that LLMs fine-tuned on WildFeedback dataset exhibit significantly improved alignment with user preferences, as evidenced by both traditional benchmarks and our proposed checklist-guided evaluation. By incorporating in-situ feedback from actual users, WildFeedback addresses the scalability, subjectivity, and bias challenges that plague existing approaches, marking a significant step toward developing LLMs that are more responsive to the diverse and evolving needs of their users.
title WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback
topic Computation and Language
url https://arxiv.org/abs/2408.15549