Beemo: Benchmark of Expert-edited Machine-generated Outputs
Fuente:
arXiv
Saved in:
| Main Authors: | Artemova, Ekaterina, Lucas, Jason, Venkatraman, Saranya, Lee, Jooyoung, Tilga, Sergei, Uchendu, Adaku, Mikhailov, Vladislav |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GPT-who: An Information Density-based Machine-Generated Text Detector
by: Venkatraman, Saranya, et al.
Published: (2023)
by: Venkatraman, Saranya, et al.
Published: (2023)
AIpom at SemEval-2024 Task 8: Detecting AI-produced Outputs in M4
by: Shirnin, Alexander, et al.
Published: (2024)
by: Shirnin, Alexander, et al.
Published: (2024)
Papilusion at DAGPap24: Paper or Illusion? Detecting AI-generated Scientific Papers
by: Andreev, Nikita, et al.
Published: (2024)
by: Andreev, Nikita, et al.
Published: (2024)
Voices of Freelance Professional Writers on AI: Limitations, Expectations, and Fears
by: Ivanova, Anastasiia, et al.
Published: (2025)
by: Ivanova, Anastasiia, et al.
Published: (2025)
DIA-HARM: Dialectal Disparities in Harmful Content Detection Across 50 English Dialects
by: Lucas, Jason, et al.
Published: (2026)
by: Lucas, Jason, et al.
Published: (2026)
RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs
by: Taktasheva, Ekaterina, et al.
Published: (2024)
by: Taktasheva, Ekaterina, et al.
Published: (2024)
Topological Data Analysis Applications in Natural Language Processing: A Survey
by: Uchendu, Adaku, et al.
Published: (2024)
by: Uchendu, Adaku, et al.
Published: (2024)
REPA: Russian Error Types Annotation for Evaluating Text Generation and Judgment Capabilities
by: Pugachev, Alexander, et al.
Published: (2025)
by: Pugachev, Alexander, et al.
Published: (2025)
PlagBench: Exploring the Duality of Large Language Models in Plagiarism Generation and Detection
by: Lee, Jooyoung, et al.
Published: (2024)
by: Lee, Jooyoung, et al.
Published: (2024)
LUNA: A Framework for Language Understanding and Naturalness Assessment
by: Saidov, Marat, et al.
Published: (2024)
by: Saidov, Marat, et al.
Published: (2024)
TOPFORMER: Topology-Aware Authorship Attribution of Deepfake Texts with Diverse Writing Styles
by: Uchendu, Adaku, et al.
Published: (2023)
by: Uchendu, Adaku, et al.
Published: (2023)
U-MATH: A University-Level Benchmark for Evaluating Mathematical Skills in LLMs
by: Chernyshev, Konstantin, et al.
Published: (2024)
by: Chernyshev, Konstantin, et al.
Published: (2024)
A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts
by: Tripto, Nafis Irtiza, et al.
Published: (2023)
by: Tripto, Nafis Irtiza, et al.
Published: (2023)
Hands-On Tutorial: Labeling with LLM and Human-in-the-Loop
by: Artemova, Ekaterina, et al.
Published: (2024)
by: Artemova, Ekaterina, et al.
Published: (2024)
Signature vs. Substance: Evaluating the Balance of Adversarial Resistance and Linguistic Quality in Watermarking Large Language Models
by: Guo, William, et al.
Published: (2025)
by: Guo, William, et al.
Published: (2025)
Beyond speculation: Measuring the growing presence of LLM-generated texts in multilingual disinformation
by: Macko, Dominik, et al.
Published: (2025)
by: Macko, Dominik, et al.
Published: (2025)
MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark
by: Macko, Dominik, et al.
Published: (2023)
by: Macko, Dominik, et al.
Published: (2023)
BLUFF: Benchmarking the Detection of False and Synthetic Content across 58 Low-Resource Languages
by: Lucas, Jason, et al.
Published: (2026)
by: Lucas, Jason, et al.
Published: (2026)
Authorship Obfuscation in Multilingual Machine-Generated Text Detection
by: Macko, Dominik, et al.
Published: (2024)
by: Macko, Dominik, et al.
Published: (2024)
Tendem: A Hybrid AI+Human Platform
by: Chernyshev, Konstantin, et al.
Published: (2026)
by: Chernyshev, Konstantin, et al.
Published: (2026)
JEEM: Vision-Language Understanding in Four Arabic Dialects
by: Kadaoui, Karima, et al.
Published: (2025)
by: Kadaoui, Karima, et al.
Published: (2025)
CollabStory: Multi-LLM Collaborative Story Generation and Authorship Analysis
by: Venkatraman, Saranya, et al.
Published: (2024)
by: Venkatraman, Saranya, et al.
Published: (2024)
Exploring the Robustness of Task-oriented Dialogue Systems for Colloquial German Varieties
by: Artemova, Ekaterina, et al.
Published: (2024)
by: Artemova, Ekaterina, et al.
Published: (2024)
Beyond checkmate: exploring the creative chokepoints in AI text
by: Tripto, Nafis Irtiza, et al.
Published: (2025)
by: Tripto, Nafis Irtiza, et al.
Published: (2025)
Masks and Mimicry: Strategic Obfuscation and Impersonation Attacks on Authorship Verification
by: Alperin, Kenneth, et al.
Published: (2025)
by: Alperin, Kenneth, et al.
Published: (2025)
RuBia: A Russian Language Bias Detection Dataset
by: Grigoreva, Veronika, et al.
Published: (2024)
by: Grigoreva, Veronika, et al.
Published: (2024)
Benchmarking Abstractive Summarisation: A Dataset of Human-authored Summaries of Norwegian News Articles
by: Touileb, Samia, et al.
Published: (2025)
by: Touileb, Samia, et al.
Published: (2025)
Donkii: Can Annotation Error Detection Methods Find Errors in Instruction-Tuning Datasets?
by: Weber-Genzel, Leon, et al.
Published: (2023)
by: Weber-Genzel, Leon, et al.
Published: (2023)
ALISON: Fast and Effective Stylometric Authorship Obfuscation
by: Xing, Eric, et al.
Published: (2024)
by: Xing, Eric, et al.
Published: (2024)
SemEval-2026 Task 4: Narrative Story Similarity and Narrative Representation Learning
by: Hatzel, Hans Ole, et al.
Published: (2026)
by: Hatzel, Hans Ole, et al.
Published: (2026)
NorEval: A Norwegian Language Understanding and Generation Evaluation Benchmark
by: Mikhailov, Vladislav, et al.
Published: (2025)
by: Mikhailov, Vladislav, et al.
Published: (2025)
Small Languages, Big Models: A Study of Continual Training on Languages of Norway
by: Samuel, David, et al.
Published: (2024)
by: Samuel, David, et al.
Published: (2024)
Sebastian, Basti, Wastl?! Recognizing Named Entities in Bavarian Dialectal Data
by: Peng, Siyao, et al.
Published: (2024)
by: Peng, Siyao, et al.
Published: (2024)
Adverbs Revisited: Enhancing WordNet Coverage of Adverbs with a Supersense Taxonomy
by: Lee, Jooyoung, et al.
Published: (2025)
by: Lee, Jooyoung, et al.
Published: (2025)
LLM-DetectAIve: a Tool for Fine-Grained Machine-Generated Text Detection
by: Abassy, Mervat, et al.
Published: (2024)
by: Abassy, Mervat, et al.
Published: (2024)
MTUncertainty: Assessing the Need for Post-editing of Machine Translation Outputs by Fine-tuning OpenAI LLMs
by: Gladkoff, Serge, et al.
Published: (2023)
by: Gladkoff, Serge, et al.
Published: (2023)
Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies
by: Artemova, Ekaterina, et al.
Published: (2025)
by: Artemova, Ekaterina, et al.
Published: (2025)
A Collection of Question Answering Datasets for Norwegian
by: Mikhailov, Vladislav, et al.
Published: (2025)
by: Mikhailov, Vladislav, et al.
Published: (2025)
DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
by: Han, Janghoon, et al.
Published: (2025)
by: Han, Janghoon, et al.
Published: (2025)
GenAI Content Detection Task 1: English and Multilingual Machine-Generated Text Detection: AI vs. Human
by: Wang, Yuxia, et al.
Published: (2025)
by: Wang, Yuxia, et al.
Published: (2025)
Similar Items
-
GPT-who: An Information Density-based Machine-Generated Text Detector
by: Venkatraman, Saranya, et al.
Published: (2023) -
AIpom at SemEval-2024 Task 8: Detecting AI-produced Outputs in M4
by: Shirnin, Alexander, et al.
Published: (2024) -
Papilusion at DAGPap24: Paper or Illusion? Detecting AI-generated Scientific Papers
by: Andreev, Nikita, et al.
Published: (2024) -
Voices of Freelance Professional Writers on AI: Limitations, Expectations, and Fears
by: Ivanova, Anastasiia, et al.
Published: (2025) -
DIA-HARM: Dialectal Disparities in Harmful Content Detection Across 50 English Dialects
by: Lucas, Jason, et al.
Published: (2026)