RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Taktasheva, Ekaterina, Bazhukov, Maxim, Koncha, Kirill, Fenogenova, Alena, Artemova, Ekaterina, Mikhailov, Vladislav |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
REPA: Russian Error Types Annotation for Evaluating Text Generation and Judgment Capabilities
von: Pugachev, Alexander, et al.
Veröffentlicht: (2025)
von: Pugachev, Alexander, et al.
Veröffentlicht: (2025)
LUNA: A Framework for Language Understanding and Naturalness Assessment
von: Saidov, Marat, et al.
Veröffentlicht: (2024)
von: Saidov, Marat, et al.
Veröffentlicht: (2024)
TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs
von: Başar, Ezgi, et al.
Veröffentlicht: (2025)
von: Başar, Ezgi, et al.
Veröffentlicht: (2025)
QFrBLiMP: a Quebec-French Benchmark of Linguistic Minimal Pairs
von: Beauchemin, David, et al.
Veröffentlicht: (2025)
von: Beauchemin, David, et al.
Veröffentlicht: (2025)
MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs
von: Jumelet, Jaap, et al.
Veröffentlicht: (2025)
von: Jumelet, Jaap, et al.
Veröffentlicht: (2025)
A Family of Pretrained Transformer Language Models for Russian
von: Zmitrovich, Dmitry, et al.
Veröffentlicht: (2023)
von: Zmitrovich, Dmitry, et al.
Veröffentlicht: (2023)
AIpom at SemEval-2024 Task 8: Detecting AI-produced Outputs in M4
von: Shirnin, Alexander, et al.
Veröffentlicht: (2024)
von: Shirnin, Alexander, et al.
Veröffentlicht: (2024)
Papilusion at DAGPap24: Paper or Illusion? Detecting AI-generated Scientific Papers
von: Andreev, Nikita, et al.
Veröffentlicht: (2024)
von: Andreev, Nikita, et al.
Veröffentlicht: (2024)
RuBia: A Russian Language Bias Detection Dataset
von: Grigoreva, Veronika, et al.
Veröffentlicht: (2024)
von: Grigoreva, Veronika, et al.
Veröffentlicht: (2024)
TAPS: Tool-Augmented Personalisation via Structured Tagging
von: Taktasheva, Ekaterina, et al.
Veröffentlicht: (2025)
von: Taktasheva, Ekaterina, et al.
Veröffentlicht: (2025)
Beemo: Benchmark of Expert-edited Machine-generated Outputs
von: Artemova, Ekaterina, et al.
Veröffentlicht: (2024)
von: Artemova, Ekaterina, et al.
Veröffentlicht: (2024)
UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu
von: Adeeba, Farah, et al.
Veröffentlicht: (2025)
von: Adeeba, Farah, et al.
Veröffentlicht: (2025)
Irish-BLiMP: A Linguistic Benchmark for Evaluating Human and Language Model Performance in a Low-Resource Setting
von: McGiff, Josh, et al.
Veröffentlicht: (2025)
von: McGiff, Josh, et al.
Veröffentlicht: (2025)
CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language Models
von: Oba, Miyu, et al.
Veröffentlicht: (2026)
von: Oba, Miyu, et al.
Veröffentlicht: (2026)
The Russian-focused embedders' exploration: ruMTEB benchmark and Russian embedding model design
von: Snegirev, Artem, et al.
Veröffentlicht: (2024)
von: Snegirev, Artem, et al.
Veröffentlicht: (2024)
Long Input Benchmark for Russian Analysis
von: Churin, Igor, et al.
Veröffentlicht: (2024)
von: Churin, Igor, et al.
Veröffentlicht: (2024)
Exploring the Robustness of Task-oriented Dialogue Systems for Colloquial German Varieties
von: Artemova, Ekaterina, et al.
Veröffentlicht: (2024)
von: Artemova, Ekaterina, et al.
Veröffentlicht: (2024)
Voices of Freelance Professional Writers on AI: Limitations, Expectations, and Fears
von: Ivanova, Anastasiia, et al.
Veröffentlicht: (2025)
von: Ivanova, Anastasiia, et al.
Veröffentlicht: (2025)
Donkii: Can Annotation Error Detection Methods Find Errors in Instruction-Tuning Datasets?
von: Weber-Genzel, Leon, et al.
Veröffentlicht: (2023)
von: Weber-Genzel, Leon, et al.
Veröffentlicht: (2023)
U-MATH: A University-Level Benchmark for Evaluating Mathematical Skills in LLMs
von: Chernyshev, Konstantin, et al.
Veröffentlicht: (2024)
von: Chernyshev, Konstantin, et al.
Veröffentlicht: (2024)
Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models
von: Zhou, Xinyu, et al.
Veröffentlicht: (2024)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2024)
SemEval-2026 Task 4: Narrative Story Similarity and Narrative Representation Learning
von: Hatzel, Hans Ole, et al.
Veröffentlicht: (2026)
von: Hatzel, Hans Ole, et al.
Veröffentlicht: (2026)
ProcrustesGPT: Compressing LLMs with Structured Matrices and Orthogonal Transformations
von: Grishina, Ekaterina, et al.
Veröffentlicht: (2025)
von: Grishina, Ekaterina, et al.
Veröffentlicht: (2025)
Eye of Judgement: Dissecting the Evaluation of Russian-speaking LLMs with POLLUX
von: Martynov, Nikita, et al.
Veröffentlicht: (2025)
von: Martynov, Nikita, et al.
Veröffentlicht: (2025)
Hands-On Tutorial: Labeling with LLM and Human-in-the-Loop
von: Artemova, Ekaterina, et al.
Veröffentlicht: (2024)
von: Artemova, Ekaterina, et al.
Veröffentlicht: (2024)
Sebastian, Basti, Wastl?! Recognizing Named Entities in Bavarian Dialectal Data
von: Peng, Siyao, et al.
Veröffentlicht: (2024)
von: Peng, Siyao, et al.
Veröffentlicht: (2024)
Benchmarking Abstractive Summarisation: A Dataset of Human-authored Summaries of Norwegian News Articles
von: Touileb, Samia, et al.
Veröffentlicht: (2025)
von: Touileb, Samia, et al.
Veröffentlicht: (2025)
Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies
von: Artemova, Ekaterina, et al.
Veröffentlicht: (2025)
von: Artemova, Ekaterina, et al.
Veröffentlicht: (2025)
A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese
von: Liu, Yikang, et al.
Veröffentlicht: (2024)
von: Liu, Yikang, et al.
Veröffentlicht: (2024)
Targeted Linguistic Analysis of Sign Language Models with Minimal Translation Pairs
von: Karabüklü, Serpil, et al.
Veröffentlicht: (2026)
von: Karabüklü, Serpil, et al.
Veröffentlicht: (2026)
Russian-Language Multimodal Dataset for Automatic Summarization of Scientific Papers
von: Tsanda, Alena, et al.
Veröffentlicht: (2024)
von: Tsanda, Alena, et al.
Veröffentlicht: (2024)
PetKaz at SemEval-2024 Task 8: Can Linguistics Capture the Specifics of LLM-generated Text?
von: Petukhova, Kseniia, et al.
Veröffentlicht: (2024)
von: Petukhova, Kseniia, et al.
Veröffentlicht: (2024)
XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs
von: He, Linyang, et al.
Veröffentlicht: (2025)
von: He, Linyang, et al.
Veröffentlicht: (2025)
Yesterday's News: Benchmarking Multi-Dimensional Out-of-Distribution Generalization of Misinformation Detection Models
von: Verhoeven, Ivo, et al.
Veröffentlicht: (2024)
von: Verhoeven, Ivo, et al.
Veröffentlicht: (2024)
PetKaz at SemEval-2024 Task 3: Advancing Emotion Classification with an LLM for Emotion-Cause Pair Extraction in Conversations
von: Kazakov, Roman, et al.
Veröffentlicht: (2024)
von: Kazakov, Roman, et al.
Veröffentlicht: (2024)
Multimodal Evaluation of Russian-language Architectures
von: Chervyakov, Artem, et al.
Veröffentlicht: (2025)
von: Chervyakov, Artem, et al.
Veröffentlicht: (2025)
MERA: A Comprehensive LLM Evaluation in Russian
von: Fenogenova, Alena, et al.
Veröffentlicht: (2024)
von: Fenogenova, Alena, et al.
Veröffentlicht: (2024)
REFeREE: A REference-FREE Model-Based Metric for Text Simplification
von: Huang, Yichen, et al.
Veröffentlicht: (2024)
von: Huang, Yichen, et al.
Veröffentlicht: (2024)
Induction Heads as an Essential Mechanism for Pattern Matching in In-context Learning
von: Crosbie, Joy, et al.
Veröffentlicht: (2024)
von: Crosbie, Joy, et al.
Veröffentlicht: (2024)
Self-Alignment: Improving Alignment of Cultural Values in LLMs via In-Context Learning
von: Choenni, Rochelle, et al.
Veröffentlicht: (2024)
von: Choenni, Rochelle, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
REPA: Russian Error Types Annotation for Evaluating Text Generation and Judgment Capabilities
von: Pugachev, Alexander, et al.
Veröffentlicht: (2025) -
LUNA: A Framework for Language Understanding and Naturalness Assessment
von: Saidov, Marat, et al.
Veröffentlicht: (2024) -
TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs
von: Başar, Ezgi, et al.
Veröffentlicht: (2025) -
QFrBLiMP: a Quebec-French Benchmark of Linguistic Minimal Pairs
von: Beauchemin, David, et al.
Veröffentlicht: (2025) -
MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs
von: Jumelet, Jaap, et al.
Veröffentlicht: (2025)