QFrBLiMP: a Quebec-French Benchmark of Linguistic Minimal Pairs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Beauchemin, David, Veilleux, Pier-Luc, Roy, Johanna-Pascale, Khoury, Richard |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
QFrCoLA: a Quebec-French Corpus of Linguistic Acceptability Judgments
von: Beauchemin, David, et al.
Veröffentlicht: (2025)
von: Beauchemin, David, et al.
Veröffentlicht: (2025)
RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs
von: Taktasheva, Ekaterina, et al.
Veröffentlicht: (2024)
von: Taktasheva, Ekaterina, et al.
Veröffentlicht: (2024)
TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs
von: Başar, Ezgi, et al.
Veröffentlicht: (2025)
von: Başar, Ezgi, et al.
Veröffentlicht: (2025)
Benchmarking Large Language Models for Quebec Insurance: From Closed-Book to Retrieval-Augmented Generation
von: Beauchemin, David, et al.
Veröffentlicht: (2026)
von: Beauchemin, David, et al.
Veröffentlicht: (2026)
MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs
von: Jumelet, Jaap, et al.
Veröffentlicht: (2025)
von: Jumelet, Jaap, et al.
Veröffentlicht: (2025)
COLE: a Comprehensive Benchmark for French Language Understanding Evaluation
von: Beauchemin, David, et al.
Veröffentlicht: (2025)
von: Beauchemin, David, et al.
Veröffentlicht: (2025)
Quebec Automobile Insurance Question-Answering With Retrieval-Augmented Generation
von: Beauchemin, David, et al.
Veröffentlicht: (2024)
von: Beauchemin, David, et al.
Veröffentlicht: (2024)
UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu
von: Adeeba, Farah, et al.
Veröffentlicht: (2025)
von: Adeeba, Farah, et al.
Veröffentlicht: (2025)
Irish-BLiMP: A Linguistic Benchmark for Evaluating Human and Language Model Performance in a Low-Resource Setting
von: McGiff, Josh, et al.
Veröffentlicht: (2025)
von: McGiff, Josh, et al.
Veröffentlicht: (2025)
CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language Models
von: Oba, Miyu, et al.
Veröffentlicht: (2026)
von: Oba, Miyu, et al.
Veröffentlicht: (2026)
JUDGEBERT: Assessing Legal Meaning Preservation Between Sentences
von: Beauchemin, David, et al.
Veröffentlicht: (2025)
von: Beauchemin, David, et al.
Veröffentlicht: (2025)
Idiom Understanding as a Tool to Measure the Dialect Gap
von: Beauchemin, David, et al.
Veröffentlicht: (2025)
von: Beauchemin, David, et al.
Veröffentlicht: (2025)
Neural Machine Translation for Coptic-French: Strategies for Low-Resource Ancient Languages
von: Chaoui, Nasma, et al.
Veröffentlicht: (2025)
von: Chaoui, Nasma, et al.
Veröffentlicht: (2025)
Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models
von: Zhou, Xinyu, et al.
Veröffentlicht: (2024)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2024)
A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese
von: Liu, Yikang, et al.
Veröffentlicht: (2024)
von: Liu, Yikang, et al.
Veröffentlicht: (2024)
Targeted Linguistic Analysis of Sign Language Models with Minimal Translation Pairs
von: Karabüklü, Serpil, et al.
Veröffentlicht: (2026)
von: Karabüklü, Serpil, et al.
Veröffentlicht: (2026)
Automated Journalistic Questions: A New Method for Extracting 5W1H in French
von: Verhaverbeke, Maxence, et al.
Veröffentlicht: (2025)
von: Verhaverbeke, Maxence, et al.
Veröffentlicht: (2025)
XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs
von: He, Linyang, et al.
Veröffentlicht: (2025)
von: He, Linyang, et al.
Veröffentlicht: (2025)
Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study
von: Khan, Eeham, et al.
Veröffentlicht: (2025)
von: Khan, Eeham, et al.
Veröffentlicht: (2025)
Decoding Probing: Revealing Internal Linguistic Structures in Neural Language Models using Minimal Pairs
von: He, Linyang, et al.
Veröffentlicht: (2024)
von: He, Linyang, et al.
Veröffentlicht: (2024)
From Rosetta to Match-Up: A Paired Corpus of Linguistic Puzzles with Human and LLM Benchmarks
von: Majmudar, Neh, et al.
Veröffentlicht: (2026)
von: Majmudar, Neh, et al.
Veröffentlicht: (2026)
BLiSS 1.0: Evaluating Bilingual Learner Competence in Second Language Small Language Models
von: Gao, Yuan, et al.
Veröffentlicht: (2025)
von: Gao, Yuan, et al.
Veröffentlicht: (2025)
FrenchToxicityPrompts: a Large Benchmark for Evaluating and Mitigating Toxicity in French Texts
von: Brun, Caroline, et al.
Veröffentlicht: (2024)
von: Brun, Caroline, et al.
Veröffentlicht: (2024)
Minimal Pair-Based Evaluation of Code-Switching
von: Sterner, Igor, et al.
Veröffentlicht: (2025)
von: Sterner, Igor, et al.
Veröffentlicht: (2025)
A Benchmark of French ASR Systems Based on Error Severity
von: Tholly, Antoine, et al.
Veröffentlicht: (2025)
von: Tholly, Antoine, et al.
Veröffentlicht: (2025)
Benchmarking Linguistic Diversity of Large Language Models
von: Guo, Yanzhu, et al.
Veröffentlicht: (2024)
von: Guo, Yanzhu, et al.
Veröffentlicht: (2024)
PragWorld: A Benchmark Evaluating LLMs' Local World Model under Minimal Linguistic Alterations and Conversational Dynamics
von: Vashistha, Sachin, et al.
Veröffentlicht: (2025)
von: Vashistha, Sachin, et al.
Veröffentlicht: (2025)
KoBALT: Korean Benchmark For Advanced Linguistic Tasks
von: Shin, Hyopil, et al.
Veröffentlicht: (2025)
von: Shin, Hyopil, et al.
Veröffentlicht: (2025)
What has LeBenchmark Learnt about French Syntax?
von: Dugonjić, Zdravko, et al.
Veröffentlicht: (2024)
von: Dugonjić, Zdravko, et al.
Veröffentlicht: (2024)
IOLBENCH: Benchmarking LLMs on Linguistic Reasoning
von: Goyal, Satyam, et al.
Veröffentlicht: (2025)
von: Goyal, Satyam, et al.
Veröffentlicht: (2025)
LongLaMP: A Benchmark for Personalized Long-form Text Generation
von: Kumar, Ishita, et al.
Veröffentlicht: (2024)
von: Kumar, Ishita, et al.
Veröffentlicht: (2024)
Discriminating Form and Meaning in Multilingual Models with Minimal-Pair ABX Tasks
von: de Seyssel, Maureen, et al.
Veröffentlicht: (2025)
von: de Seyssel, Maureen, et al.
Veröffentlicht: (2025)
PBEBench: A Multi-Step Programming by Examples Reasoning Benchmark inspired by Historical Linguistics
von: Naik, Atharva, et al.
Veröffentlicht: (2025)
von: Naik, Atharva, et al.
Veröffentlicht: (2025)
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models
von: Waldis, Andreas, et al.
Veröffentlicht: (2024)
von: Waldis, Andreas, et al.
Veröffentlicht: (2024)
Exploring Gaps in the APS: Direct Minimal Pair Analysis in LLM Syntactic Assessments
von: Pistotti, Timothy, et al.
Veröffentlicht: (2025)
von: Pistotti, Timothy, et al.
Veröffentlicht: (2025)
DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
SPOT: An Annotated French Corpus and Benchmark for Detecting Critical Interventions in Online Conversations
von: Berriche, Manon, et al.
Veröffentlicht: (2025)
von: Berriche, Manon, et al.
Veröffentlicht: (2025)
How Linguistics Learned to Stop Worrying and Love the Language Models
von: Futrell, Richard, et al.
Veröffentlicht: (2025)
von: Futrell, Richard, et al.
Veröffentlicht: (2025)
LaMP-QA: A Benchmark for Personalized Long-form Question Answering
von: Salemi, Alireza, et al.
Veröffentlicht: (2025)
von: Salemi, Alireza, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
QFrCoLA: a Quebec-French Corpus of Linguistic Acceptability Judgments
von: Beauchemin, David, et al.
Veröffentlicht: (2025) -
RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs
von: Taktasheva, Ekaterina, et al.
Veröffentlicht: (2024) -
TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs
von: Başar, Ezgi, et al.
Veröffentlicht: (2025) -
Benchmarking Large Language Models for Quebec Insurance: From Closed-Book to Retrieval-Augmented Generation
von: Beauchemin, David, et al.
Veröffentlicht: (2026) -
MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs
von: Jumelet, Jaap, et al.
Veröffentlicht: (2025)