Diversity Explains Inference Scaling Laws: Through a Case Study of Minimum Bayes Risk Decoding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kamigaito, Hidetaka, Deguchi, Hiroyuki, Sakai, Yusuke, Hayashi, Katsuhiko, Watanabe, Taro
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913882195361792
author Kamigaito, Hidetaka
Deguchi, Hiroyuki
Sakai, Yusuke
Hayashi, Katsuhiko
Watanabe, Taro
author_facet Kamigaito, Hidetaka
Deguchi, Hiroyuki
Sakai, Yusuke
Hayashi, Katsuhiko
Watanabe, Taro
contents Inference methods play an important role in eliciting the performance of large language models (LLMs). Currently, LLMs use inference methods utilizing generated multiple samples, which can be derived from Minimum Bayes Risk (MBR) Decoding. Previous studies have conducted empirical analyses to clarify the improvements in generation performance achieved by MBR decoding and have reported various observations. However, the theoretical underpinnings of these findings remain uncertain. To address this, we offer a new theoretical interpretation of MBR decoding from the perspective of bias-diversity decomposition. In this interpretation, the error in the quality estimation of hypotheses by MBR decoding is decomposed into two main factors: bias, which considers the closeness between the utility function and human evaluation, and diversity, which represents the variability in the quality estimation of the utility function. The theoretical analysis reveals the difficulty of simultaneously improving bias and diversity, confirming the validity of enhancing MBR decoding performance by increasing diversity. Furthermore, we reveal that diversity can explain one aspect of inference scaling laws that describe performance improvement by increasing sample size. Moreover, experiments across multiple NLP tasks yielded results consistent with these theoretical characteristics. Our code is available at https://github.com/naist-nlp/mbr-bias-diversity.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15021
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Diversity Explains Inference Scaling Laws: Through a Case Study of Minimum Bayes Risk Decoding
Kamigaito, Hidetaka
Deguchi, Hiroyuki
Sakai, Yusuke
Hayashi, Katsuhiko
Watanabe, Taro
Computation and Language
Inference methods play an important role in eliciting the performance of large language models (LLMs). Currently, LLMs use inference methods utilizing generated multiple samples, which can be derived from Minimum Bayes Risk (MBR) Decoding. Previous studies have conducted empirical analyses to clarify the improvements in generation performance achieved by MBR decoding and have reported various observations. However, the theoretical underpinnings of these findings remain uncertain. To address this, we offer a new theoretical interpretation of MBR decoding from the perspective of bias-diversity decomposition. In this interpretation, the error in the quality estimation of hypotheses by MBR decoding is decomposed into two main factors: bias, which considers the closeness between the utility function and human evaluation, and diversity, which represents the variability in the quality estimation of the utility function. The theoretical analysis reveals the difficulty of simultaneously improving bias and diversity, confirming the validity of enhancing MBR decoding performance by increasing diversity. Furthermore, we reveal that diversity can explain one aspect of inference scaling laws that describe performance improvement by increasing sample size. Moreover, experiments across multiple NLP tasks yielded results consistent with these theoretical characteristics. Our code is available at https://github.com/naist-nlp/mbr-bias-diversity.
title Diversity Explains Inference Scaling Laws: Through a Case Study of Minimum Bayes Risk Decoding
topic Computation and Language
url https://arxiv.org/abs/2410.15021