MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs
Fuente:
arXiv
Saved in:
| Main Authors: | Jumelet, Jaap, Weissweiler, Leonie, Nivre, Joakim, Bisazza, Arianna |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs
by: Başar, Ezgi, et al.
Published: (2025)
by: Başar, Ezgi, et al.
Published: (2025)
RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs
by: Taktasheva, Ekaterina, et al.
Published: (2024)
by: Taktasheva, Ekaterina, et al.
Published: (2024)
Assessing the Impact of Typological Features on Multilingual Machine Translation in the Age of Large Language Models
by: Hirak, Vitalii, et al.
Published: (2026)
by: Hirak, Vitalii, et al.
Published: (2026)
QFrBLiMP: a Quebec-French Benchmark of Linguistic Minimal Pairs
by: Beauchemin, David, et al.
Published: (2025)
by: Beauchemin, David, et al.
Published: (2025)
Is Child-Directed Language Optimized for Word Learning? A Computational Study of Verb Meaning Acquisition
by: Padovani, Francesca, et al.
Published: (2026)
by: Padovani, Francesca, et al.
Published: (2026)
Child-Directed Language Does Not Consistently Boost Syntax Learning in Language Models
by: Padovani, Francesca, et al.
Published: (2025)
by: Padovani, Francesca, et al.
Published: (2025)
UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu
by: Adeeba, Farah, et al.
Published: (2025)
by: Adeeba, Farah, et al.
Published: (2025)
Linguistic Generalizations are not Rules: Impacts on Evaluation of LMs
by: Weissweiler, Leonie, et al.
Published: (2025)
by: Weissweiler, Leonie, et al.
Published: (2025)
Hybrid Human-LLM Corpus Construction and LLM Evaluation for Rare Linguistic Phenomena
by: Weissweiler, Leonie, et al.
Published: (2024)
by: Weissweiler, Leonie, et al.
Published: (2024)
CAIT: A Syntactic Parsing Toolkit for Child-Adult InTeractions
by: Padovani, Francesca, et al.
Published: (2026)
by: Padovani, Francesca, et al.
Published: (2026)
Irish-BLiMP: A Linguistic Benchmark for Evaluating Human and Language Model Performance in a Low-Resource Setting
by: McGiff, Josh, et al.
Published: (2025)
by: McGiff, Josh, et al.
Published: (2025)
Finding Structure in Language Models
by: Jumelet, Jaap
Published: (2024)
by: Jumelet, Jaap
Published: (2024)
CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language Models
by: Oba, Miyu, et al.
Published: (2026)
by: Oba, Miyu, et al.
Published: (2026)
Climate-Eval: A Comprehensive Benchmark for NLP Tasks Related to Climate Change
by: Kurfalı, Murathan, et al.
Published: (2025)
by: Kurfalı, Murathan, et al.
Published: (2025)
Cross-Lingual Transfer of Debiasing and Detoxification in Multilingual LLMs: An Extensive Investigation
by: Neplenbroek, Vera, et al.
Published: (2024)
by: Neplenbroek, Vera, et al.
Published: (2024)
On the Consistency of Multilingual Context Utilization in Retrieval-Augmented Generation
by: Qi, Jirui, et al.
Published: (2025)
by: Qi, Jirui, et al.
Published: (2025)
Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models
by: Qi, Jirui, et al.
Published: (2023)
by: Qi, Jirui, et al.
Published: (2023)
BabyLM's First Constructions: Causal probing provides a signal of learning
by: Rozner, Joshua, et al.
Published: (2025)
by: Rozner, Joshua, et al.
Published: (2025)
XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs
by: He, Linyang, et al.
Published: (2025)
by: He, Linyang, et al.
Published: (2025)
Language Models Learn Constructional Semantics, Not To Mention Syntax: Investigating LM Understanding of Paired-Focus Constructions
by: Scivetti, Wesley, et al.
Published: (2026)
by: Scivetti, Wesley, et al.
Published: (2026)
Do Language Models Exhibit Human-like Structural Priming Effects?
by: Jumelet, Jaap, et al.
Published: (2024)
by: Jumelet, Jaap, et al.
Published: (2024)
Continual Learning Under Language Shift
by: Gogoulou, Evangelia, et al.
Published: (2023)
by: Gogoulou, Evangelia, et al.
Published: (2023)
MBBQ: A Dataset for Cross-Lingual Comparison of Stereotypes in Generative LLMs
by: Neplenbroek, Vera, et al.
Published: (2024)
by: Neplenbroek, Vera, et al.
Published: (2024)
NeLLCom-X: A Comprehensive Neural-Agent Framework to Simulate Language Learning and Group Communication
by: Lian, Yuchen, et al.
Published: (2024)
by: Lian, Yuchen, et al.
Published: (2024)
UCxn: Typologically Informed Annotation of Constructions Atop Universal Dependencies
by: Weissweiler, Leonie, et al.
Published: (2024)
by: Weissweiler, Leonie, et al.
Published: (2024)
BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data
by: Jumelet, Jaap, et al.
Published: (2025)
by: Jumelet, Jaap, et al.
Published: (2025)
Constructions are Revealed in Word Distributions
by: Rozner, Joshua, et al.
Published: (2025)
by: Rozner, Joshua, et al.
Published: (2025)
Both Direct and Indirect Evidence Contribute to Dative Alternation Preferences in Language Models
by: Yao, Qing, et al.
Published: (2025)
by: Yao, Qing, et al.
Published: (2025)
BLiSS 1.0: Evaluating Bilingual Learner Competence in Second Language Small Language Models
by: Gao, Yuan, et al.
Published: (2025)
by: Gao, Yuan, et al.
Published: (2025)
Reading Between the Prompts: How Stereotypes Shape LLM's Implicit Personalization
by: Neplenbroek, Vera, et al.
Published: (2025)
by: Neplenbroek, Vera, et al.
Published: (2025)
Simulating the Emergence of Differential Case Marking with Communicating Neural-Network Agents
by: Lian, Yuchen, et al.
Published: (2025)
by: Lian, Yuchen, et al.
Published: (2025)
Black Big Boxes: Tracing Adjective Order Preferences in Large Language Models
by: Jumelet, Jaap, et al.
Published: (2024)
by: Jumelet, Jaap, et al.
Published: (2024)
Models Can and Should Embrace the Communicative Nature of Human-Generated Math
by: Boguraev, Sasha, et al.
Published: (2024)
by: Boguraev, Sasha, et al.
Published: (2024)
Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models
by: Zhou, Xinyu, et al.
Published: (2024)
by: Zhou, Xinyu, et al.
Published: (2024)
Can LLMs Detect Intrinsic Hallucinations in Paraphrasing and Machine Translation?
by: Gogoulou, Evangelia, et al.
Published: (2025)
by: Gogoulou, Evangelia, et al.
Published: (2025)
The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text Generation
by: Carlsson, Fredrik, et al.
Published: (2024)
by: Carlsson, Fredrik, et al.
Published: (2024)
MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval
by: Kriz, Reno, et al.
Published: (2024)
by: Kriz, Reno, et al.
Published: (2024)
Interpretability of Language Models via Task Spaces
by: Weber, Lucas, et al.
Published: (2024)
by: Weber, Lucas, et al.
Published: (2024)
A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese
by: Liu, Yikang, et al.
Published: (2024)
by: Liu, Yikang, et al.
Published: (2024)
Verbing Weirds Language (Models): Evaluation of English Zero-Derivation in Five LLMs
by: Mortensen, David R., et al.
Published: (2024)
by: Mortensen, David R., et al.
Published: (2024)
Similar Items
-
TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs
by: Başar, Ezgi, et al.
Published: (2025) -
RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs
by: Taktasheva, Ekaterina, et al.
Published: (2024) -
Assessing the Impact of Typological Features on Multilingual Machine Translation in the Age of Large Language Models
by: Hirak, Vitalii, et al.
Published: (2026) -
QFrBLiMP: a Quebec-French Benchmark of Linguistic Minimal Pairs
by: Beauchemin, David, et al.
Published: (2025) -
Is Child-Directed Language Optimized for Word Learning? A Computational Study of Verb Meaning Acquisition
by: Padovani, Francesca, et al.
Published: (2026)