MetaAlign: Align Large Language Models with Diverse Preferences during Inference Time

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Mozhi, Wang, Pengyu, Tan, Chenkun, Huang, Mianqiu, Zhang, Dong, Zhou, Yaqian, Qiu, Xipeng
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929549752664064
author Zhang, Mozhi
Wang, Pengyu
Tan, Chenkun
Huang, Mianqiu
Zhang, Dong
Zhou, Yaqian
Qiu, Xipeng
author_facet Zhang, Mozhi
Wang, Pengyu
Tan, Chenkun
Huang, Mianqiu
Zhang, Dong
Zhou, Yaqian
Qiu, Xipeng
contents Large Language Models (LLMs) acquire extensive knowledge and remarkable abilities from extensive text corpora, making them powerful tools for various applications. To make LLMs more usable, aligning them with human preferences is essential. Existing alignment techniques, such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO), typically embed predefined preferences directly within the model's parameters. These methods, however, often result in a static alignment that can not account for the diversity of human preferences in practical applications. In response to this challenge, we propose an effective method, \textbf{MetaAlign}, which aims to help LLMs dynamically align with various explicit or implicit preferences specified at inference time. Experimental results show that LLMs optimized on our meticulously constructed MetaAlign Dataset can effectively align with any preferences specified at the inference stage, validating the feasibility of MetaAlign. We hope that our work can provide some insights into the alignment of language models.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14184
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MetaAlign: Align Large Language Models with Diverse Preferences during Inference Time
Zhang, Mozhi
Wang, Pengyu
Tan, Chenkun
Huang, Mianqiu
Zhang, Dong
Zhou, Yaqian
Qiu, Xipeng
Computation and Language
Large Language Models (LLMs) acquire extensive knowledge and remarkable abilities from extensive text corpora, making them powerful tools for various applications. To make LLMs more usable, aligning them with human preferences is essential. Existing alignment techniques, such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO), typically embed predefined preferences directly within the model's parameters. These methods, however, often result in a static alignment that can not account for the diversity of human preferences in practical applications. In response to this challenge, we propose an effective method, \textbf{MetaAlign}, which aims to help LLMs dynamically align with various explicit or implicit preferences specified at inference time. Experimental results show that LLMs optimized on our meticulously constructed MetaAlign Dataset can effectively align with any preferences specified at the inference stage, validating the feasibility of MetaAlign. We hope that our work can provide some insights into the alignment of language models.
title MetaAlign: Align Large Language Models with Diverse Preferences during Inference Time
topic Computation and Language
url https://arxiv.org/abs/2410.14184