XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912250006077440 |
|---|---|
| author | He, Linyang Nie, Ercong Dindar, Sukru Samet Firoozi, Arsalan Florea, Adrian Nguyen, Van Puffay, Corentin Shimizu, Riki Ye, Haotian Brennan, Jonathan Schmid, Helmut Schütze, Hinrich Mesgarani, Nima |
| author_facet | He, Linyang Nie, Ercong Dindar, Sukru Samet Firoozi, Arsalan Florea, Adrian Nguyen, Van Puffay, Corentin Shimizu, Riki Ye, Haotian Brennan, Jonathan Schmid, Helmut Schütze, Hinrich Mesgarani, Nima |
| contents | We introduce XCOMPS in this work, a multilingual conceptual minimal pair dataset covering 17 languages. Using this dataset, we evaluate LLMs' multilingual conceptual understanding through metalinguistic prompting, direct probability measurement, and neurolinguistic probing. By comparing base, instruction-tuned, and knowledge-distilled models, we find that: 1) LLMs exhibit weaker conceptual understanding for low-resource languages, and accuracy varies across languages despite being tested on the same concept sets. 2) LLMs excel at distinguishing concept-property pairs that are visibly different but exhibit a marked performance drop when negative pairs share subtle semantic similarities. 3) Instruction tuning improves performance in concept understanding but does not enhance internal competence; knowledge distillation can enhance internal competence in conceptual understanding for low-resource languages with limited gains in explicit task performance. 4) More morphologically complex languages yield lower concept understanding scores and require deeper layers for conceptual reasoning. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_19737 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs He, Linyang Nie, Ercong Dindar, Sukru Samet Firoozi, Arsalan Florea, Adrian Nguyen, Van Puffay, Corentin Shimizu, Riki Ye, Haotian Brennan, Jonathan Schmid, Helmut Schütze, Hinrich Mesgarani, Nima Computation and Language We introduce XCOMPS in this work, a multilingual conceptual minimal pair dataset covering 17 languages. Using this dataset, we evaluate LLMs' multilingual conceptual understanding through metalinguistic prompting, direct probability measurement, and neurolinguistic probing. By comparing base, instruction-tuned, and knowledge-distilled models, we find that: 1) LLMs exhibit weaker conceptual understanding for low-resource languages, and accuracy varies across languages despite being tested on the same concept sets. 2) LLMs excel at distinguishing concept-property pairs that are visibly different but exhibit a marked performance drop when negative pairs share subtle semantic similarities. 3) Instruction tuning improves performance in concept understanding but does not enhance internal competence; knowledge distillation can enhance internal competence in conceptual understanding for low-resource languages with limited gains in explicit task performance. 4) More morphologically complex languages yield lower concept understanding scores and require deeper layers for conceptual reasoning. |
| title | XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2502.19737 |