M-IFEval: Multilingual Instruction-Following Evaluation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Dussolle, Antoine, Díaz, Andrea Cardeña, Sato, Shota, Devine, Peter
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917916173139968
author Dussolle, Antoine
Díaz, Andrea Cardeña
Sato, Shota
Devine, Peter
author_facet Dussolle, Antoine
Díaz, Andrea Cardeña
Sato, Shota
Devine, Peter
contents Instruction following is a core capability of modern Large language models (LLMs), making evaluating this capability essential to understanding these models. The Instruction Following Evaluation (IFEval) benchmark from the literature does this using objective criteria, offering a measure of LLM performance without subjective AI or human judgement. However, it only includes English instructions, limiting its ability to assess LLMs in other languages. We propose the Multilingual Instruction Following Evaluation (M-IFEval) benchmark, expanding the evaluation to French, Japanese, and Spanish, with both general and language-specific instructions. Applying this benchmark to 8 state-of-the-art LLMs, we find that benchmark performance across languages and instruction types can vary widely, underscoring the importance of a multilingual benchmark for evaluating LLMs in a diverse cultural context.
format Preprint
id arxiv_https___arxiv_org_abs_2502_04688
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle M-IFEval: Multilingual Instruction-Following Evaluation
Dussolle, Antoine
Díaz, Andrea Cardeña
Sato, Shota
Devine, Peter
Computation and Language
Artificial Intelligence
Instruction following is a core capability of modern Large language models (LLMs), making evaluating this capability essential to understanding these models. The Instruction Following Evaluation (IFEval) benchmark from the literature does this using objective criteria, offering a measure of LLM performance without subjective AI or human judgement. However, it only includes English instructions, limiting its ability to assess LLMs in other languages. We propose the Multilingual Instruction Following Evaluation (M-IFEval) benchmark, expanding the evaluation to French, Japanese, and Spanish, with both general and language-specific instructions. Applying this benchmark to 8 state-of-the-art LLMs, we find that benchmark performance across languages and instruction types can vary widely, underscoring the importance of a multilingual benchmark for evaluating LLMs in a diverse cultural context.
title M-IFEval: Multilingual Instruction-Following Evaluation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.04688