Prompting with Phonemes: Enhancing LLMs' Multilinguality for Non-Latin Script Languages

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Nguyen, Hoang H, Mahajan, Khyati, Yadav, Vikas, Salazar, Julian, Yu, Philip S., Hashemi, Masoud, Maheshwary, Rishabh
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909660421816320
author Nguyen, Hoang H
Mahajan, Khyati
Yadav, Vikas
Salazar, Julian
Yu, Philip S.
Hashemi, Masoud
Maheshwary, Rishabh
author_facet Nguyen, Hoang H
Mahajan, Khyati
Yadav, Vikas
Salazar, Julian
Yu, Philip S.
Hashemi, Masoud
Maheshwary, Rishabh
contents Although multilingual LLMs have achieved remarkable performance across benchmarks, we find they continue to underperform on non-Latin script languages across contemporary LLM families. This discrepancy arises from the fact that LLMs are pretrained with orthographic scripts, which are dominated by Latin characters that obscure their shared phonology with non-Latin scripts. We propose leveraging phonemic transcriptions as complementary signals to induce script-invariant representations. Our study demonstrates that integrating phonemic signals improves performance across both non-Latin and Latin script languages, with a particularly significant impact on closing the performance gap between the two. Through detailed experiments, we show that phonemic and orthographic scripts retrieve distinct examples for in-context learning (ICL). This motivates our proposed Mixed-ICL retrieval strategy, where further aggregation from both leads to our significant performance improvements for both Latin script languages (up to 12.6%) and non-Latin script languages (up to 15.1%) compared to randomized ICL retrieval.
format Preprint
id arxiv_https___arxiv_org_abs_2411_02398
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Prompting with Phonemes: Enhancing LLMs' Multilinguality for Non-Latin Script Languages
Nguyen, Hoang H
Mahajan, Khyati
Yadav, Vikas
Salazar, Julian
Yu, Philip S.
Hashemi, Masoud
Maheshwary, Rishabh
Computation and Language
Artificial Intelligence
Machine Learning
Although multilingual LLMs have achieved remarkable performance across benchmarks, we find they continue to underperform on non-Latin script languages across contemporary LLM families. This discrepancy arises from the fact that LLMs are pretrained with orthographic scripts, which are dominated by Latin characters that obscure their shared phonology with non-Latin scripts. We propose leveraging phonemic transcriptions as complementary signals to induce script-invariant representations. Our study demonstrates that integrating phonemic signals improves performance across both non-Latin and Latin script languages, with a particularly significant impact on closing the performance gap between the two. Through detailed experiments, we show that phonemic and orthographic scripts retrieve distinct examples for in-context learning (ICL). This motivates our proposed Mixed-ICL retrieval strategy, where further aggregation from both leads to our significant performance improvements for both Latin script languages (up to 12.6%) and non-Latin script languages (up to 15.1%) compared to randomized ICL retrieval.
title Prompting with Phonemes: Enhancing LLMs' Multilinguality for Non-Latin Script Languages
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.02398