Open Conversational LLMs do not know most Spanish words

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Conde, Javier, González, Miguel, Melero, Nina, Ferrando, Raquel, Martínez, Gonzalo, Merino-Gómez, Elena, Hernández, José Alberto, Reviriego, Pedro
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913513785524224
author Conde, Javier
González, Miguel
Melero, Nina
Ferrando, Raquel
Martínez, Gonzalo
Merino-Gómez, Elena
Hernández, José Alberto
Reviriego, Pedro
author_facet Conde, Javier
González, Miguel
Melero, Nina
Ferrando, Raquel
Martínez, Gonzalo
Merino-Gómez, Elena
Hernández, José Alberto
Reviriego, Pedro
contents The growing interest in Large Language Models (LLMs) and in particular in conversational models with which users can interact has led to the development of a large number of open-source chat LLMs. These models are evaluated on a wide range of benchmarks to assess their capabilities in answering questions or solving problems on almost any possible topic or to test their ability to reason or interpret texts. Instead, the evaluation of the knowledge that these models have of the languages has received much less attention. For example, the words that they can recognize and use in different languages. In this paper, we evaluate the knowledge that open-source chat LLMs have of Spanish words by testing a sample of words in a reference dictionary. The results show that open-source chat LLMs produce incorrect meanings for an important fraction of the words and are not able to use most of the words correctly to write sentences with context. These results show how Spanish is left behind in the open-source LLM race and highlight the need to push for linguistic fairness in conversational LLMs ensuring that they provide similar performance across languages.
format Preprint
id arxiv_https___arxiv_org_abs_2403_15491
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Open Conversational LLMs do not know most Spanish words
Conde, Javier
González, Miguel
Melero, Nina
Ferrando, Raquel
Martínez, Gonzalo
Merino-Gómez, Elena
Hernández, José Alberto
Reviriego, Pedro
Computation and Language
The growing interest in Large Language Models (LLMs) and in particular in conversational models with which users can interact has led to the development of a large number of open-source chat LLMs. These models are evaluated on a wide range of benchmarks to assess their capabilities in answering questions or solving problems on almost any possible topic or to test their ability to reason or interpret texts. Instead, the evaluation of the knowledge that these models have of the languages has received much less attention. For example, the words that they can recognize and use in different languages. In this paper, we evaluate the knowledge that open-source chat LLMs have of Spanish words by testing a sample of words in a reference dictionary. The results show that open-source chat LLMs produce incorrect meanings for an important fraction of the words and are not able to use most of the words correctly to write sentences with context. These results show how Spanish is left behind in the open-source LLM race and highlight the need to push for linguistic fairness in conversational LLMs ensuring that they provide similar performance across languages.
title Open Conversational LLMs do not know most Spanish words
topic Computation and Language
url https://arxiv.org/abs/2403.15491