CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Oba, Miyu, Sugawara, Saku |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Language Models Induce Grammatical Knowledge from Indirect Evidence?
by: Oba, Miyu, et al.
Published: (2024)
by: Oba, Miyu, et al.
Published: (2024)
RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs
by: Taktasheva, Ekaterina, et al.
Published: (2024)
by: Taktasheva, Ekaterina, et al.
Published: (2024)
TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs
by: Başar, Ezgi, et al.
Published: (2025)
by: Başar, Ezgi, et al.
Published: (2025)
A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language Models
by: Emura, Rei, et al.
Published: (2026)
by: Emura, Rei, et al.
Published: (2026)
What Makes Language Models Good-enough?
by: Asami, Daiki, et al.
Published: (2024)
by: Asami, Daiki, et al.
Published: (2024)
QFrBLiMP: a Quebec-French Benchmark of Linguistic Minimal Pairs
by: Beauchemin, David, et al.
Published: (2025)
by: Beauchemin, David, et al.
Published: (2025)
Rationale-Aware Answer Verification by Pairwise Self-Evaluation
by: Kawabata, Akira, et al.
Published: (2024)
by: Kawabata, Akira, et al.
Published: (2024)
Specification-Aware Machine Translation and Evaluation for Purpose Alignment
by: Kayano, Yoko, et al.
Published: (2025)
by: Kayano, Yoko, et al.
Published: (2025)
MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs
by: Jumelet, Jaap, et al.
Published: (2025)
by: Jumelet, Jaap, et al.
Published: (2025)
C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences
by: Kawabata, Akira, et al.
Published: (2026)
by: Kawabata, Akira, et al.
Published: (2026)
UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu
by: Adeeba, Farah, et al.
Published: (2025)
by: Adeeba, Farah, et al.
Published: (2025)
Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models
by: Zhou, Xinyu, et al.
Published: (2024)
by: Zhou, Xinyu, et al.
Published: (2024)
Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?
by: Furuhashi, Momoka, et al.
Published: (2025)
by: Furuhashi, Momoka, et al.
Published: (2025)
A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese
by: Liu, Yikang, et al.
Published: (2024)
by: Liu, Yikang, et al.
Published: (2024)
Irish-BLiMP: A Linguistic Benchmark for Evaluating Human and Language Model Performance in a Low-Resource Setting
by: McGiff, Josh, et al.
Published: (2025)
by: McGiff, Josh, et al.
Published: (2025)
Targeted Linguistic Analysis of Sign Language Models with Minimal Translation Pairs
by: Karabüklü, Serpil, et al.
Published: (2026)
by: Karabüklü, Serpil, et al.
Published: (2026)
TactfulToM: Do LLMs Have the Theory of Mind Ability to Understand White Lies?
by: Liu, Yiwei, et al.
Published: (2025)
by: Liu, Yiwei, et al.
Published: (2025)
BabyLM Challenge: Exploring the Effect of Variation Sets on Language Model Training Efficiency
by: Haga, Akari, et al.
Published: (2024)
by: Haga, Akari, et al.
Published: (2024)
BQA: Body Language Question Answering Dataset for Video Large Language Models
by: Ozaki, Shintaro, et al.
Published: (2024)
by: Ozaki, Shintaro, et al.
Published: (2024)
Evaluating CxG Generalisation in LLMs via Construction-Based NLI Fine Tuning
by: Mackintosh, Tom, et al.
Published: (2025)
by: Mackintosh, Tom, et al.
Published: (2025)
Decoding Probing: Revealing Internal Linguistic Structures in Neural Language Models using Minimal Pairs
by: He, Linyang, et al.
Published: (2024)
by: He, Linyang, et al.
Published: (2024)
XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs
by: He, Linyang, et al.
Published: (2025)
by: He, Linyang, et al.
Published: (2025)
What Matters in Memorizing and Recalling Facts? Multifaceted Benchmarks for Knowledge Probing in Language Models
by: Zhao, Xin, et al.
Published: (2024)
by: Zhao, Xin, et al.
Published: (2024)
Which Feedback Works for Whom? Differential Effects of LLM-Generated Feedback Elements Across Learner Profiles
by: Furuhashi, Momoka, et al.
Published: (2026)
by: Furuhashi, Momoka, et al.
Published: (2026)
MoreHopQA: More Than Multi-hop Reasoning
by: Schnitzler, Julian, et al.
Published: (2024)
by: Schnitzler, Julian, et al.
Published: (2024)
Benchmarking Linguistic Diversity of Large Language Models
by: Guo, Yanzhu, et al.
Published: (2024)
by: Guo, Yanzhu, et al.
Published: (2024)
Language Models Learn Constructional Semantics, Not To Mention Syntax: Investigating LM Understanding of Paired-Focus Constructions
by: Scivetti, Wesley, et al.
Published: (2026)
by: Scivetti, Wesley, et al.
Published: (2026)
Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models
by: Waseda, Futa, et al.
Published: (2025)
by: Waseda, Futa, et al.
Published: (2025)
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models
by: Waldis, Andreas, et al.
Published: (2024)
by: Waldis, Andreas, et al.
Published: (2024)
Minimal Pair-Based Evaluation of Code-Switching
by: Sterner, Igor, et al.
Published: (2025)
by: Sterner, Igor, et al.
Published: (2025)
PragWorld: A Benchmark Evaluating LLMs' Local World Model under Minimal Linguistic Alterations and Conversational Dynamics
by: Vashistha, Sachin, et al.
Published: (2025)
by: Vashistha, Sachin, et al.
Published: (2025)
From Rosetta to Match-Up: A Paired Corpus of Linguistic Puzzles with Human and LLM Benchmarks
by: Majmudar, Neh, et al.
Published: (2026)
by: Majmudar, Neh, et al.
Published: (2026)
TLUE: A Tibetan Language Understanding Evaluation Benchmark
by: Gao, Fan, et al.
Published: (2025)
by: Gao, Fan, et al.
Published: (2025)
VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages
by: Atuhurra, Jesse, et al.
Published: (2025)
by: Atuhurra, Jesse, et al.
Published: (2025)
Measuring Human Involvement in AI-Generated Text: A Case Study on Academic Writing
by: Guo, Yuchen, et al.
Published: (2025)
by: Guo, Yuchen, et al.
Published: (2025)
Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge
by: Sun, Xin, et al.
Published: (2026)
by: Sun, Xin, et al.
Published: (2026)
Automatic Feedback Generation for Short Answer Questions using Answer Diagnostic Graphs
by: Furuhashi, Momoka, et al.
Published: (2025)
by: Furuhashi, Momoka, et al.
Published: (2025)
BLUCK: A Benchmark Dataset for Bengali Linguistic Understanding and Cultural Knowledge
by: Kabir, Daeen, et al.
Published: (2025)
by: Kabir, Daeen, et al.
Published: (2025)
Evaluating List Construction and Temporal Understanding capabilities of Large Language Models
by: Dumitru, Alexandru, et al.
Published: (2025)
by: Dumitru, Alexandru, et al.
Published: (2025)
The Invalsi Benchmarks: measuring Linguistic and Mathematical understanding of Large Language Models in Italian
by: Puccetti, Giovanni, et al.
Published: (2024)
by: Puccetti, Giovanni, et al.
Published: (2024)
Similar Items
-
Can Language Models Induce Grammatical Knowledge from Indirect Evidence?
by: Oba, Miyu, et al.
Published: (2024) -
RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs
by: Taktasheva, Ekaterina, et al.
Published: (2024) -
TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs
by: Başar, Ezgi, et al.
Published: (2025) -
A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language Models
by: Emura, Rei, et al.
Published: (2026) -
What Makes Language Models Good-enough?
by: Asami, Daiki, et al.
Published: (2024)