How Does Quantization Affect Multilingual LLMs?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Marchisio, Kelly, Dash, Saurabh, Chen, Hongyu, Aumiller, Dennis, Üstün, Ahmet, Hooker, Sara, Ruder, Sebastian
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929538409168896
author Marchisio, Kelly
Dash, Saurabh
Chen, Hongyu
Aumiller, Dennis
Üstün, Ahmet
Hooker, Sara
Ruder, Sebastian
author_facet Marchisio, Kelly
Dash, Saurabh
Chen, Hongyu
Aumiller, Dennis
Üstün, Ahmet
Hooker, Sara
Ruder, Sebastian
contents Quantization techniques are widely used to improve inference speed and deployment of large language models. While a wide body of work examines the impact of quantization on LLMs in English, none have evaluated across languages. We conduct a thorough analysis of quantized multilingual LLMs, focusing on performance across languages and at varying scales. We use automatic benchmarks, LLM-as-a-Judge, and human evaluation, finding that (1) harmful effects of quantization are apparent in human evaluation, which automatic metrics severely underestimate: a 1.7% average drop in Japanese across automatic tasks corresponds to a 16.0% drop reported by human evaluators on realistic prompts; (2) languages are disparately affected by quantization, with non-Latin script languages impacted worst; and (3) challenging tasks like mathematical reasoning degrade fastest. As the ability to serve low-compute models is critical for wide global adoption of NLP technologies, our results urge consideration of multilingual performance as a key evaluation criterion for efficient models.
format Preprint
id arxiv_https___arxiv_org_abs_2407_03211
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How Does Quantization Affect Multilingual LLMs?
Marchisio, Kelly
Dash, Saurabh
Chen, Hongyu
Aumiller, Dennis
Üstün, Ahmet
Hooker, Sara
Ruder, Sebastian
Computation and Language
Machine Learning
Quantization techniques are widely used to improve inference speed and deployment of large language models. While a wide body of work examines the impact of quantization on LLMs in English, none have evaluated across languages. We conduct a thorough analysis of quantized multilingual LLMs, focusing on performance across languages and at varying scales. We use automatic benchmarks, LLM-as-a-Judge, and human evaluation, finding that (1) harmful effects of quantization are apparent in human evaluation, which automatic metrics severely underestimate: a 1.7% average drop in Japanese across automatic tasks corresponds to a 16.0% drop reported by human evaluators on realistic prompts; (2) languages are disparately affected by quantization, with non-Latin script languages impacted worst; and (3) challenging tasks like mathematical reasoning degrade fastest. As the ability to serve low-compute models is critical for wide global adoption of NLP technologies, our results urge consideration of multilingual performance as a key evaluation criterion for efficient models.
title How Does Quantization Affect Multilingual LLMs?
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2407.03211