A Probability--Quality Trade-off in Aligned Language Models and its Relation to Sampling Adaptors
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866912089178636288 |
|---|---|
| author | Tan, Naaman Valvoda, Josef Liu, Tianyu Svete, Anej Qin, Yanxia Min-Yen, Kan Cotterell, Ryan |
| author_facet | Tan, Naaman Valvoda, Josef Liu, Tianyu Svete, Anej Qin, Yanxia Min-Yen, Kan Cotterell, Ryan |
| contents | The relationship between the quality of a string, as judged by a human reader, and its probability, $p(\boldsymbol{y})$ under a language model undergirds the development of better language models. For example, many popular algorithms for sampling from a language model have been conceived with the goal of manipulating $p(\boldsymbol{y})$ to place higher probability on strings that humans deem of high quality. In this article, we examine the probability--quality relationship in language models explicitly aligned to human preferences, e.g., through reinforcement learning through human feedback. We show that, when sampling corpora from an aligned language model, there exists a trade-off between the strings' average reward and average log-likelihood under the prior language model, i.e., the same model before alignment with human preferences. We provide a formal treatment of this phenomenon and demonstrate how a choice of sampling adaptor allows for a selection of how much likelihood we exchange for the reward. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_10203 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | A Probability--Quality Trade-off in Aligned Language Models and its Relation to Sampling Adaptors Tan, Naaman Valvoda, Josef Liu, Tianyu Svete, Anej Qin, Yanxia Min-Yen, Kan Cotterell, Ryan Computation and Language The relationship between the quality of a string, as judged by a human reader, and its probability, $p(\boldsymbol{y})$ under a language model undergirds the development of better language models. For example, many popular algorithms for sampling from a language model have been conceived with the goal of manipulating $p(\boldsymbol{y})$ to place higher probability on strings that humans deem of high quality. In this article, we examine the probability--quality relationship in language models explicitly aligned to human preferences, e.g., through reinforcement learning through human feedback. We show that, when sampling corpora from an aligned language model, there exists a trade-off between the strings' average reward and average log-likelihood under the prior language model, i.e., the same model before alignment with human preferences. We provide a formal treatment of this phenomenon and demonstrate how a choice of sampling adaptor allows for a selection of how much likelihood we exchange for the reward. |
| title | A Probability--Quality Trade-off in Aligned Language Models and its Relation to Sampling Adaptors |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2406.10203 |