Since the Scientific Literature Is Multilingual, Our Models Should Be Too

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ebrahimi, Abteen, Church, Kenneth
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916179752255488
author Ebrahimi, Abteen
Church, Kenneth
author_facet Ebrahimi, Abteen
Church, Kenneth
contents English has long been assumed the $\textit{lingua franca}$ of scientific research, and this notion is reflected in the natural language processing (NLP) research involving scientific document representation. In this position piece, we quantitatively show that the literature is largely multilingual and argue that current models and benchmarks should reflect this linguistic diversity. We provide evidence that text-based models fail to create meaningful representations for non-English papers and highlight the negative user-facing impacts of using English-only models non-discriminately across a multilingual domain. We end with suggestions for the NLP community on how to improve performance on non-English documents.
format Preprint
id arxiv_https___arxiv_org_abs_2403_18251
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Since the Scientific Literature Is Multilingual, Our Models Should Be Too
Ebrahimi, Abteen
Church, Kenneth
Computation and Language
English has long been assumed the $\textit{lingua franca}$ of scientific research, and this notion is reflected in the natural language processing (NLP) research involving scientific document representation. In this position piece, we quantitatively show that the literature is largely multilingual and argue that current models and benchmarks should reflect this linguistic diversity. We provide evidence that text-based models fail to create meaningful representations for non-English papers and highlight the negative user-facing impacts of using English-only models non-discriminately across a multilingual domain. We end with suggestions for the NLP community on how to improve performance on non-English documents.
title Since the Scientific Literature Is Multilingual, Our Models Should Be Too
topic Computation and Language
url https://arxiv.org/abs/2403.18251