MaXIFE: Multilingual and Cross-lingual Instruction Following Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yile, Ma, Ziwei, Jiang, Xiu, Hu, Jinglu, Chang, Jing, Li, Liang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916774198378496
author Liu, Yile
Ma, Ziwei
Jiang, Xiu
Hu, Jinglu
Chang, Jing
Li, Liang
author_facet Liu, Yile
Ma, Ziwei
Jiang, Xiu
Hu, Jinglu
Chang, Jing
Li, Liang
contents With the rapid adoption of large language models (LLMs) in natural language processing, the ability to follow instructions has emerged as a key metric for evaluating their practical utility. However, existing evaluation methods often focus on single-language scenarios, overlooking the challenges and differences present in multilingual and cross-lingual contexts. To address this gap, we introduce MaXIFE: a comprehensive evaluation benchmark designed to assess instruction-following capabilities across 23 different languages with 1667 verifiable instruction tasks. MaXIFE integrates both Rule-Based Evaluation and Model-Based Evaluation, ensuring a balance of efficiency and accuracy. We applied MaXIFE to evaluate several leading commercial LLMs, establishing baseline results for future comparisons. By providing a standardized tool for multilingual instruction-following evaluation, MaXIFE aims to advance research and development in natural language processing.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01776
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MaXIFE: Multilingual and Cross-lingual Instruction Following Evaluation
Liu, Yile
Ma, Ziwei
Jiang, Xiu
Hu, Jinglu
Chang, Jing
Li, Liang
Computation and Language
Artificial Intelligence
With the rapid adoption of large language models (LLMs) in natural language processing, the ability to follow instructions has emerged as a key metric for evaluating their practical utility. However, existing evaluation methods often focus on single-language scenarios, overlooking the challenges and differences present in multilingual and cross-lingual contexts. To address this gap, we introduce MaXIFE: a comprehensive evaluation benchmark designed to assess instruction-following capabilities across 23 different languages with 1667 verifiable instruction tasks. MaXIFE integrates both Rule-Based Evaluation and Model-Based Evaluation, ensuring a balance of efficiency and accuracy. We applied MaXIFE to evaluate several leading commercial LLMs, establishing baseline results for future comparisons. By providing a standardized tool for multilingual instruction-following evaluation, MaXIFE aims to advance research and development in natural language processing.
title MaXIFE: Multilingual and Cross-lingual Instruction Following Evaluation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.01776