A Probability--Quality Trade-off in Aligned Language Models and its Relation to Sampling Adaptors

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Tan, Naaman, Valvoda, Josef, Liu, Tianyu, Svete, Anej, Qin, Yanxia, Min-Yen, Kan, Cotterell, Ryan
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912089178636288
author Tan, Naaman
Valvoda, Josef
Liu, Tianyu
Svete, Anej
Qin, Yanxia
Min-Yen, Kan
Cotterell, Ryan
author_facet Tan, Naaman
Valvoda, Josef
Liu, Tianyu
Svete, Anej
Qin, Yanxia
Min-Yen, Kan
Cotterell, Ryan
contents The relationship between the quality of a string, as judged by a human reader, and its probability, $p(\boldsymbol{y})$ under a language model undergirds the development of better language models. For example, many popular algorithms for sampling from a language model have been conceived with the goal of manipulating $p(\boldsymbol{y})$ to place higher probability on strings that humans deem of high quality. In this article, we examine the probability--quality relationship in language models explicitly aligned to human preferences, e.g., through reinforcement learning through human feedback. We show that, when sampling corpora from an aligned language model, there exists a trade-off between the strings' average reward and average log-likelihood under the prior language model, i.e., the same model before alignment with human preferences. We provide a formal treatment of this phenomenon and demonstrate how a choice of sampling adaptor allows for a selection of how much likelihood we exchange for the reward.
format Preprint
id arxiv_https___arxiv_org_abs_2406_10203
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Probability--Quality Trade-off in Aligned Language Models and its Relation to Sampling Adaptors
Tan, Naaman
Valvoda, Josef
Liu, Tianyu
Svete, Anej
Qin, Yanxia
Min-Yen, Kan
Cotterell, Ryan
Computation and Language
The relationship between the quality of a string, as judged by a human reader, and its probability, $p(\boldsymbol{y})$ under a language model undergirds the development of better language models. For example, many popular algorithms for sampling from a language model have been conceived with the goal of manipulating $p(\boldsymbol{y})$ to place higher probability on strings that humans deem of high quality. In this article, we examine the probability--quality relationship in language models explicitly aligned to human preferences, e.g., through reinforcement learning through human feedback. We show that, when sampling corpora from an aligned language model, there exists a trade-off between the strings' average reward and average log-likelihood under the prior language model, i.e., the same model before alignment with human preferences. We provide a formal treatment of this phenomenon and demonstrate how a choice of sampling adaptor allows for a selection of how much likelihood we exchange for the reward.
title A Probability--Quality Trade-off in Aligned Language Models and its Relation to Sampling Adaptors
topic Computation and Language
url https://arxiv.org/abs/2406.10203