Aligning Large Language Models with Implicit Preferences from User-Generated Content

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Zhaoxuan, Li, Zheng, Liu, Tianyi, Wang, Haodong, Yun, Hyokun, Zeng, Ming, Chen, Pei, Zhang, Zhihan, Gao, Yifan, Wang, Ruijie, Nigam, Priyanka, Yin, Bing, Jiang, Meng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909638841073664
author Tan, Zhaoxuan
Li, Zheng
Liu, Tianyi
Wang, Haodong
Yun, Hyokun
Zeng, Ming
Chen, Pei
Zhang, Zhihan
Gao, Yifan
Wang, Ruijie
Nigam, Priyanka
Yin, Bing
Jiang, Meng
author_facet Tan, Zhaoxuan
Li, Zheng
Liu, Tianyi
Wang, Haodong
Yun, Hyokun
Zeng, Ming
Chen, Pei
Zhang, Zhihan
Gao, Yifan
Wang, Ruijie
Nigam, Priyanka
Yin, Bing
Jiang, Meng
contents Learning from preference feedback is essential for aligning large language models (LLMs) with human values and improving the quality of generated responses. However, existing preference learning methods rely heavily on curated data from humans or advanced LLMs, which is costly and difficult to scale. In this work, we present PUGC, a novel framework that leverages implicit human Preferences in unlabeled User-Generated Content (UGC) to generate preference data. Although UGC is not explicitly created to guide LLMs in generating human-preferred responses, it often reflects valuable insights and implicit preferences from its creators that has the potential to address readers' questions. PUGC transforms UGC into user queries and generates responses from the policy model. The UGC is then leveraged as a reference text for response scoring, aligning the model with these implicit preferences. This approach improves the quality of preference data while enabling scalable, domain-specific alignment. Experimental results on Alpaca Eval 2 show that models trained with DPO and PUGC achieve a 9.37% performance improvement over traditional methods, setting a 35.93% state-of-the-art length-controlled win rate using Mistral-7B-Instruct. Further studies highlight gains in reward quality, domain-specific alignment effectiveness, robustness against UGC quality, and theory of mind capabilities. Our code and dataset are available at https://zhaoxuan.info/PUGC.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2506_04463
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Aligning Large Language Models with Implicit Preferences from User-Generated Content
Tan, Zhaoxuan
Li, Zheng
Liu, Tianyi
Wang, Haodong
Yun, Hyokun
Zeng, Ming
Chen, Pei
Zhang, Zhihan
Gao, Yifan
Wang, Ruijie
Nigam, Priyanka
Yin, Bing
Jiang, Meng
Computation and Language
Learning from preference feedback is essential for aligning large language models (LLMs) with human values and improving the quality of generated responses. However, existing preference learning methods rely heavily on curated data from humans or advanced LLMs, which is costly and difficult to scale. In this work, we present PUGC, a novel framework that leverages implicit human Preferences in unlabeled User-Generated Content (UGC) to generate preference data. Although UGC is not explicitly created to guide LLMs in generating human-preferred responses, it often reflects valuable insights and implicit preferences from its creators that has the potential to address readers' questions. PUGC transforms UGC into user queries and generates responses from the policy model. The UGC is then leveraged as a reference text for response scoring, aligning the model with these implicit preferences. This approach improves the quality of preference data while enabling scalable, domain-specific alignment. Experimental results on Alpaca Eval 2 show that models trained with DPO and PUGC achieve a 9.37% performance improvement over traditional methods, setting a 35.93% state-of-the-art length-controlled win rate using Mistral-7B-Instruct. Further studies highlight gains in reward quality, domain-specific alignment effectiveness, robustness against UGC quality, and theory of mind capabilities. Our code and dataset are available at https://zhaoxuan.info/PUGC.github.io/
title Aligning Large Language Models with Implicit Preferences from User-Generated Content
topic Computation and Language
url https://arxiv.org/abs/2506.04463