Establishing Vocabulary Tests as a Benchmark for Evaluating Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Martínez, Gonzalo, Conde, Javier, Merino-Gómez, Elena, Bermúdez-Margaretto, Beatriz, Hernández, José Alberto, Reviriego, Pedro, Brysbaert, Marc
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929225896820736
author Martínez, Gonzalo
Conde, Javier
Merino-Gómez, Elena
Bermúdez-Margaretto, Beatriz
Hernández, José Alberto
Reviriego, Pedro
Brysbaert, Marc
author_facet Martínez, Gonzalo
Conde, Javier
Merino-Gómez, Elena
Bermúdez-Margaretto, Beatriz
Hernández, José Alberto
Reviriego, Pedro
Brysbaert, Marc
contents Vocabulary tests, once a cornerstone of language modeling evaluation, have been largely overlooked in the current landscape of Large Language Models (LLMs) like Llama, Mistral, and GPT. While most LLM evaluation benchmarks focus on specific tasks or domain-specific knowledge, they often neglect the fundamental linguistic aspects of language understanding and production. In this paper, we advocate for the revival of vocabulary tests as a valuable tool for assessing LLM performance. We evaluate seven LLMs using two vocabulary test formats across two languages and uncover surprising gaps in their lexical knowledge. These findings shed light on the intricacies of LLM word representations, their learning mechanisms, and performance variations across models and languages. Moreover, the ability to automatically generate and perform vocabulary tests offers new opportunities to expand the approach and provide a more complete picture of LLMs' language skills.
format Preprint
id arxiv_https___arxiv_org_abs_2310_14703
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Establishing Vocabulary Tests as a Benchmark for Evaluating Large Language Models
Martínez, Gonzalo
Conde, Javier
Merino-Gómez, Elena
Bermúdez-Margaretto, Beatriz
Hernández, José Alberto
Reviriego, Pedro
Brysbaert, Marc
Computation and Language
Vocabulary tests, once a cornerstone of language modeling evaluation, have been largely overlooked in the current landscape of Large Language Models (LLMs) like Llama, Mistral, and GPT. While most LLM evaluation benchmarks focus on specific tasks or domain-specific knowledge, they often neglect the fundamental linguistic aspects of language understanding and production. In this paper, we advocate for the revival of vocabulary tests as a valuable tool for assessing LLM performance. We evaluate seven LLMs using two vocabulary test formats across two languages and uncover surprising gaps in their lexical knowledge. These findings shed light on the intricacies of LLM word representations, their learning mechanisms, and performance variations across models and languages. Moreover, the ability to automatically generate and perform vocabulary tests offers new opportunities to expand the approach and provide a more complete picture of LLMs' language skills.
title Establishing Vocabulary Tests as a Benchmark for Evaluating Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2310.14703