XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Linyang, Nie, Ercong, Dindar, Sukru Samet, Firoozi, Arsalan, Florea, Adrian, Nguyen, Van, Puffay, Corentin, Shimizu, Riki, Ye, Haotian, Brennan, Jonathan, Schmid, Helmut, Schütze, Hinrich, Mesgarani, Nima
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912250006077440
author He, Linyang
Nie, Ercong
Dindar, Sukru Samet
Firoozi, Arsalan
Florea, Adrian
Nguyen, Van
Puffay, Corentin
Shimizu, Riki
Ye, Haotian
Brennan, Jonathan
Schmid, Helmut
Schütze, Hinrich
Mesgarani, Nima
author_facet He, Linyang
Nie, Ercong
Dindar, Sukru Samet
Firoozi, Arsalan
Florea, Adrian
Nguyen, Van
Puffay, Corentin
Shimizu, Riki
Ye, Haotian
Brennan, Jonathan
Schmid, Helmut
Schütze, Hinrich
Mesgarani, Nima
contents We introduce XCOMPS in this work, a multilingual conceptual minimal pair dataset covering 17 languages. Using this dataset, we evaluate LLMs' multilingual conceptual understanding through metalinguistic prompting, direct probability measurement, and neurolinguistic probing. By comparing base, instruction-tuned, and knowledge-distilled models, we find that: 1) LLMs exhibit weaker conceptual understanding for low-resource languages, and accuracy varies across languages despite being tested on the same concept sets. 2) LLMs excel at distinguishing concept-property pairs that are visibly different but exhibit a marked performance drop when negative pairs share subtle semantic similarities. 3) Instruction tuning improves performance in concept understanding but does not enhance internal competence; knowledge distillation can enhance internal competence in conceptual understanding for low-resource languages with limited gains in explicit task performance. 4) More morphologically complex languages yield lower concept understanding scores and require deeper layers for conceptual reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2502_19737
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs
He, Linyang
Nie, Ercong
Dindar, Sukru Samet
Firoozi, Arsalan
Florea, Adrian
Nguyen, Van
Puffay, Corentin
Shimizu, Riki
Ye, Haotian
Brennan, Jonathan
Schmid, Helmut
Schütze, Hinrich
Mesgarani, Nima
Computation and Language
We introduce XCOMPS in this work, a multilingual conceptual minimal pair dataset covering 17 languages. Using this dataset, we evaluate LLMs' multilingual conceptual understanding through metalinguistic prompting, direct probability measurement, and neurolinguistic probing. By comparing base, instruction-tuned, and knowledge-distilled models, we find that: 1) LLMs exhibit weaker conceptual understanding for low-resource languages, and accuracy varies across languages despite being tested on the same concept sets. 2) LLMs excel at distinguishing concept-property pairs that are visibly different but exhibit a marked performance drop when negative pairs share subtle semantic similarities. 3) Instruction tuning improves performance in concept understanding but does not enhance internal competence; knowledge distillation can enhance internal competence in conceptual understanding for low-resource languages with limited gains in explicit task performance. 4) More morphologically complex languages yield lower concept understanding scores and require deeper layers for conceptual reasoning.
title XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs
topic Computation and Language
url https://arxiv.org/abs/2502.19737