TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Başar, Ezgi, Padovani, Francesca, Jumelet, Jaap, Bisazza, Arianna
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911297421967360
author Başar, Ezgi
Padovani, Francesca
Jumelet, Jaap
Bisazza, Arianna
author_facet Başar, Ezgi
Padovani, Francesca
Jumelet, Jaap
Bisazza, Arianna
contents We introduce TurBLiMP, the first Turkish benchmark of linguistic minimal pairs, designed to evaluate the linguistic abilities of monolingual and multilingual language models (LMs). Covering 16 linguistic phenomena with 1000 minimal pairs each, TurBLiMP fills an important gap in linguistic evaluation resources for Turkish. In designing the benchmark, we give extra attention to two properties of Turkish that remain understudied in current syntactic evaluations of LMs, namely word order flexibility and subordination through morphological processes. Our experiments on a wide range of LMs and a newly collected set of human acceptability judgments reveal that even cutting-edge Large LMs still struggle with grammatical phenomena that are not challenging for humans, and may also exhibit different sensitivities to word order and morphological complexity compared to humans.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13487
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs
Başar, Ezgi
Padovani, Francesca
Jumelet, Jaap
Bisazza, Arianna
Computation and Language
We introduce TurBLiMP, the first Turkish benchmark of linguistic minimal pairs, designed to evaluate the linguistic abilities of monolingual and multilingual language models (LMs). Covering 16 linguistic phenomena with 1000 minimal pairs each, TurBLiMP fills an important gap in linguistic evaluation resources for Turkish. In designing the benchmark, we give extra attention to two properties of Turkish that remain understudied in current syntactic evaluations of LMs, namely word order flexibility and subordination through morphological processes. Our experiments on a wide range of LMs and a newly collected set of human acceptability judgments reveal that even cutting-edge Large LMs still struggle with grammatical phenomena that are not challenging for humans, and may also exhibit different sensitivities to word order and morphological complexity compared to humans.
title TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs
topic Computation and Language
url https://arxiv.org/abs/2506.13487