REPA: Russian Error Types Annotation for Evaluating Text Generation and Judgment Capabilities
Fuente:
arXiv
Saved in:
| Main Authors: | Pugachev, Alexander, Fenogenova, Alena, Mikhailov, Vladislav, Artemova, Ekaterina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs
by: Taktasheva, Ekaterina, et al.
Published: (2024)
by: Taktasheva, Ekaterina, et al.
Published: (2024)
AIpom at SemEval-2024 Task 8: Detecting AI-produced Outputs in M4
by: Shirnin, Alexander, et al.
Published: (2024)
by: Shirnin, Alexander, et al.
Published: (2024)
Papilusion at DAGPap24: Paper or Illusion? Detecting AI-generated Scientific Papers
by: Andreev, Nikita, et al.
Published: (2024)
by: Andreev, Nikita, et al.
Published: (2024)
LUNA: A Framework for Language Understanding and Naturalness Assessment
by: Saidov, Marat, et al.
Published: (2024)
by: Saidov, Marat, et al.
Published: (2024)
A Family of Pretrained Transformer Language Models for Russian
by: Zmitrovich, Dmitry, et al.
Published: (2023)
by: Zmitrovich, Dmitry, et al.
Published: (2023)
The Russian-focused embedders' exploration: ruMTEB benchmark and Russian embedding model design
by: Snegirev, Artem, et al.
Published: (2024)
by: Snegirev, Artem, et al.
Published: (2024)
Donkii: Can Annotation Error Detection Methods Find Errors in Instruction-Tuning Datasets?
by: Weber-Genzel, Leon, et al.
Published: (2023)
by: Weber-Genzel, Leon, et al.
Published: (2023)
Beemo: Benchmark of Expert-edited Machine-generated Outputs
by: Artemova, Ekaterina, et al.
Published: (2024)
by: Artemova, Ekaterina, et al.
Published: (2024)
RuBia: A Russian Language Bias Detection Dataset
by: Grigoreva, Veronika, et al.
Published: (2024)
by: Grigoreva, Veronika, et al.
Published: (2024)
Long Input Benchmark for Russian Analysis
by: Churin, Igor, et al.
Published: (2024)
by: Churin, Igor, et al.
Published: (2024)
Eye of Judgement: Dissecting the Evaluation of Russian-speaking LLMs with POLLUX
by: Martynov, Nikita, et al.
Published: (2025)
by: Martynov, Nikita, et al.
Published: (2025)
Multimodal Evaluation of Russian-language Architectures
by: Chervyakov, Artem, et al.
Published: (2025)
by: Chervyakov, Artem, et al.
Published: (2025)
Exploring the Robustness of Task-oriented Dialogue Systems for Colloquial German Varieties
by: Artemova, Ekaterina, et al.
Published: (2024)
by: Artemova, Ekaterina, et al.
Published: (2024)
MERA: A Comprehensive LLM Evaluation in Russian
by: Fenogenova, Alena, et al.
Published: (2024)
by: Fenogenova, Alena, et al.
Published: (2024)
RILEC: Detection and Generation of L1 Russian Interference Errors in English Learner Texts
by: Kharlamova, Darya, et al.
Published: (2026)
by: Kharlamova, Darya, et al.
Published: (2026)
Voices of Freelance Professional Writers on AI: Limitations, Expectations, and Fears
by: Ivanova, Anastasiia, et al.
Published: (2025)
by: Ivanova, Anastasiia, et al.
Published: (2025)
Russian-Language Multimodal Dataset for Automatic Summarization of Scientific Papers
by: Tsanda, Alena, et al.
Published: (2024)
by: Tsanda, Alena, et al.
Published: (2024)
LLM-DetectAIve: a Tool for Fine-Grained Machine-Generated Text Detection
by: Abassy, Mervat, et al.
Published: (2024)
by: Abassy, Mervat, et al.
Published: (2024)
From Text to Emotion: Unveiling the Emotion Annotation Capabilities of LLMs
by: Niu, Minxue, et al.
Published: (2024)
by: Niu, Minxue, et al.
Published: (2024)
NorEval: A Norwegian Language Understanding and Generation Evaluation Benchmark
by: Mikhailov, Vladislav, et al.
Published: (2025)
by: Mikhailov, Vladislav, et al.
Published: (2025)
Remedy-R: Generative Reasoning for Machine Translation Evaluation without Error Annotations
by: Tan, Shaomu, et al.
Published: (2025)
by: Tan, Shaomu, et al.
Published: (2025)
U-MATH: A University-Level Benchmark for Evaluating Mathematical Skills in LLMs
by: Chernyshev, Konstantin, et al.
Published: (2024)
by: Chernyshev, Konstantin, et al.
Published: (2024)
DRAGOn: Designing RAG On Periodically Updated Corpus
by: Chernogorskii, Fedor, et al.
Published: (2025)
by: Chernogorskii, Fedor, et al.
Published: (2025)
SemEval-2026 Task 4: Narrative Story Similarity and Narrative Representation Learning
by: Hatzel, Hans Ole, et al.
Published: (2026)
by: Hatzel, Hans Ole, et al.
Published: (2026)
Benchmarking Abstractive Summarisation: A Dataset of Human-authored Summaries of Norwegian News Articles
by: Touileb, Samia, et al.
Published: (2025)
by: Touileb, Samia, et al.
Published: (2025)
A Fully Automated Pipeline for Conversational Discourse Annotation: Tree Scheme Generation and Labeling with Large Language Models
by: Petukhova, Kseniia, et al.
Published: (2025)
by: Petukhova, Kseniia, et al.
Published: (2025)
Tracing How Annotators Think: Augmenting Preference Judgments with Reading Processes
by: de Langis, Karin, et al.
Published: (2025)
by: de Langis, Karin, et al.
Published: (2025)
HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models
by: Que, Haoran, et al.
Published: (2024)
by: Que, Haoran, et al.
Published: (2024)
SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks
by: Adamenko, Pavel, et al.
Published: (2025)
by: Adamenko, Pavel, et al.
Published: (2025)
GenAI Content Detection Task 1: English and Multilingual Machine-Generated Text Detection: AI vs. Human
by: Wang, Yuxia, et al.
Published: (2025)
by: Wang, Yuxia, et al.
Published: (2025)
Evaluating Knowledge Generation and Self-Refinement Strategies for LLM-based Column Type Annotation
by: Korini, Keti, et al.
Published: (2025)
by: Korini, Keti, et al.
Published: (2025)
Hands-On Tutorial: Labeling with LLM and Human-in-the-Loop
by: Artemova, Ekaterina, et al.
Published: (2024)
by: Artemova, Ekaterina, et al.
Published: (2024)
Sebastian, Basti, Wastl?! Recognizing Named Entities in Bavarian Dialectal Data
by: Peng, Siyao, et al.
Published: (2024)
by: Peng, Siyao, et al.
Published: (2024)
RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation
by: Vasilev, Viacheslav, et al.
Published: (2025)
by: Vasilev, Viacheslav, et al.
Published: (2025)
Hidden Measurement Error in LLM Pipelines Distorts Annotation, Evaluation, and Benchmarking
by: Messing, Solomon
Published: (2026)
by: Messing, Solomon
Published: (2026)
Summarisation of German Judgments in conjunction with a Class-based Evaluation
by: Steffes, Bianca, et al.
Published: (2025)
by: Steffes, Bianca, et al.
Published: (2025)
Intent Matters: Enhancing AI Tutoring with Fine-Grained Pedagogical Intent Annotation
by: Petukhova, Kseniia, et al.
Published: (2025)
by: Petukhova, Kseniia, et al.
Published: (2025)
Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies
by: Artemova, Ekaterina, et al.
Published: (2025)
by: Artemova, Ekaterina, et al.
Published: (2025)
Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation
by: Kocmi, Tom, et al.
Published: (2024)
by: Kocmi, Tom, et al.
Published: (2024)
AR-BENCH: Benchmarking Legal Reasoning with Judgment Error Detection, Classification and Correction
by: Li, Yifei, et al.
Published: (2026)
by: Li, Yifei, et al.
Published: (2026)
Similar Items
-
RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs
by: Taktasheva, Ekaterina, et al.
Published: (2024) -
AIpom at SemEval-2024 Task 8: Detecting AI-produced Outputs in M4
by: Shirnin, Alexander, et al.
Published: (2024) -
Papilusion at DAGPap24: Paper or Illusion? Detecting AI-generated Scientific Papers
by: Andreev, Nikita, et al.
Published: (2024) -
LUNA: A Framework for Language Understanding and Naturalness Assessment
by: Saidov, Marat, et al.
Published: (2024) -
A Family of Pretrained Transformer Language Models for Russian
by: Zmitrovich, Dmitry, et al.
Published: (2023)