Value Alignment from Unstructured Text

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Padhi, Inkit, Ramamurthy, Karthikeyan Natesan, Sattigeri, Prasanna, Nagireddy, Manish, Dognin, Pierre, Varshney, Kush R.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916362681581568
author Padhi, Inkit
Ramamurthy, Karthikeyan Natesan
Sattigeri, Prasanna
Nagireddy, Manish
Dognin, Pierre
Varshney, Kush R.
author_facet Padhi, Inkit
Ramamurthy, Karthikeyan Natesan
Sattigeri, Prasanna
Nagireddy, Manish
Dognin, Pierre
Varshney, Kush R.
contents Aligning large language models (LLMs) to value systems has emerged as a significant area of research within the fields of AI and NLP. Currently, this alignment process relies on the availability of high-quality supervised and preference data, which can be both time-consuming and expensive to curate or annotate. In this paper, we introduce a systematic end-to-end methodology for aligning LLMs to the implicit and explicit values represented in unstructured text data. Our proposed approach leverages the use of scalable synthetic data generation techniques to effectively align the model to the values present in the unstructured data. Through two distinct use-cases, we demonstrate the efficiency of our methodology on the Mistral-7B-Instruct model. Our approach credibly aligns LLMs to the values embedded within documents, and shows improved performance against other approaches, as quantified through the use of automatic metrics and win rates.
format Preprint
id arxiv_https___arxiv_org_abs_2408_10392
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Value Alignment from Unstructured Text
Padhi, Inkit
Ramamurthy, Karthikeyan Natesan
Sattigeri, Prasanna
Nagireddy, Manish
Dognin, Pierre
Varshney, Kush R.
Computation and Language
Machine Learning
Aligning large language models (LLMs) to value systems has emerged as a significant area of research within the fields of AI and NLP. Currently, this alignment process relies on the availability of high-quality supervised and preference data, which can be both time-consuming and expensive to curate or annotate. In this paper, we introduce a systematic end-to-end methodology for aligning LLMs to the implicit and explicit values represented in unstructured text data. Our proposed approach leverages the use of scalable synthetic data generation techniques to effectively align the model to the values present in the unstructured data. Through two distinct use-cases, we demonstrate the efficiency of our methodology on the Mistral-7B-Instruct model. Our approach credibly aligns LLMs to the values embedded within documents, and shows improved performance against other approaches, as quantified through the use of automatic metrics and win rates.
title Value Alignment from Unstructured Text
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2408.10392