Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kunilovskaya, Maria, Bhatia, Gagan, Albertelli, Lisa Sophie, Chen, Yanran, Greisinger, Christian, Kiefer, Lotta, Leiter, Christoph, Roy, Subhadeep, Achamaleh, Tewodros, Manzoor, Muhammad Arslan, Pohl, Sebastian, Hou, Yufang, Eger, Steffen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics
von: Roy, Subhadeep, et al.
Veröffentlicht: (2026)
von: Roy, Subhadeep, et al.
Veröffentlicht: (2026)
TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning
von: Greisinger, Christian, et al.
Veröffentlicht: (2026)
von: Greisinger, Christian, et al.
Veröffentlicht: (2026)
GerAV: Towards New Heights in German Authorship Verification using Fine-Tuned LLMs on a New Benchmark
von: Kiefer, Lotta, et al.
Veröffentlicht: (2026)
von: Kiefer, Lotta, et al.
Veröffentlicht: (2026)
DeepSeek-R1 vs. o3-mini: How Well can Reasoning LLMs Evaluate MT and Summarization?
von: Larionov, Daniil, et al.
Veröffentlicht: (2025)
von: Larionov, Daniil, et al.
Veröffentlicht: (2025)
PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation
von: Leiter, Christoph, et al.
Veröffentlicht: (2024)
von: Leiter, Christoph, et al.
Veröffentlicht: (2024)
Do Emotions Really Affect Argument Convincingness? A Dynamic Approach with LLM-based Manipulation Checks
von: Chen, Yanran, et al.
Veröffentlicht: (2025)
von: Chen, Yanran, et al.
Veröffentlicht: (2025)
BMX: Boosting Natural Language Generation Metrics with Explainability
von: Leiter, Christoph, et al.
Veröffentlicht: (2022)
von: Leiter, Christoph, et al.
Veröffentlicht: (2022)
Is there really a Citation Age Bias in NLP?
von: Nguyen, Hoa, et al.
Veröffentlicht: (2024)
von: Nguyen, Hoa, et al.
Veröffentlicht: (2024)
NLLG Quarterly arXiv Report 09/24: What are the most influential current AI Papers?
von: Leiter, Christoph, et al.
Veröffentlicht: (2024)
von: Leiter, Christoph, et al.
Veröffentlicht: (2024)
CROC: Evaluating and Training T2I Metrics with Pseudo- and Human-Labeled Contrastive Robustness Checks
von: Leiter, Christoph, et al.
Veröffentlicht: (2025)
von: Leiter, Christoph, et al.
Veröffentlicht: (2025)
Evaluating Diversity in Automatic Poetry Generation
von: Chen, Yanran, et al.
Veröffentlicht: (2024)
von: Chen, Yanran, et al.
Veröffentlicht: (2024)
Translationese as a Rational Response to Translation Task Difficulty
von: Kunilovskaya, Maria
Veröffentlicht: (2026)
von: Kunilovskaya, Maria
Veröffentlicht: (2026)
Emotionally Charged, Logically Blurred: AI-driven Emotional Framing Impairs Human Fallacy Detection
von: Chen, Yanran, et al.
Veröffentlicht: (2025)
von: Chen, Yanran, et al.
Veröffentlicht: (2025)
Towards Explainable Evaluation Metrics for Machine Translation
von: Leiter, Christoph, et al.
Veröffentlicht: (2023)
von: Leiter, Christoph, et al.
Veröffentlicht: (2023)
EPIC-EuroParl-UdS: Information-Theoretic Perspectives on Translation and Interpreting
von: Kunilovskaya, Maria, et al.
Veröffentlicht: (2026)
von: Kunilovskaya, Maria, et al.
Veröffentlicht: (2026)
Who and What? Using Linguistic Features and Annotator Characteristics to Analyze Annotation Variation
von: Maurer, Maximilian, et al.
Veröffentlicht: (2026)
von: Maurer, Maximilian, et al.
Veröffentlicht: (2026)
Syntactic Language Change in English and German: Metrics, Parsers, and Convergences
von: Chen, Yanran, et al.
Veröffentlicht: (2024)
von: Chen, Yanran, et al.
Veröffentlicht: (2024)
Annotator-Centric Active Learning for Subjective NLP Tasks
von: van der Meer, Michiel, et al.
Veröffentlicht: (2024)
von: van der Meer, Michiel, et al.
Veröffentlicht: (2024)
ValueGround: Evaluating Culture-Conditioned Visual Value Grounding in MLLMs
von: Wang, Zhipin, et al.
Veröffentlicht: (2026)
von: Wang, Zhipin, et al.
Veröffentlicht: (2026)
ByGPT5: End-to-End Style-conditioned Poetry Generation with Token-free Language Models
von: Belouadi, Jonas, et al.
Veröffentlicht: (2022)
von: Belouadi, Jonas, et al.
Veröffentlicht: (2022)
PromptOptMe: Error-Aware Prompt Compression for LLM-based MT Evaluation Metrics
von: Larionov, Daniil, et al.
Veröffentlicht: (2024)
von: Larionov, Daniil, et al.
Veröffentlicht: (2024)
USCORE: An Effective Approach to Fully Unsupervised Evaluation Metrics for Machine Translation
von: Belouadi, Jonas, et al.
Veröffentlicht: (2022)
von: Belouadi, Jonas, et al.
Veröffentlicht: (2022)
LLM-based multi-agent poetry generation in non-cooperative environments
von: Zhang, Ran, et al.
Veröffentlicht: (2024)
von: Zhang, Ran, et al.
Veröffentlicht: (2024)
BatchGEMBA: Token-Efficient Machine Translation Evaluation with Batched Prompting and Prompt Compression
von: Larionov, Daniil, et al.
Veröffentlicht: (2025)
von: Larionov, Daniil, et al.
Veröffentlicht: (2025)
Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images
von: Mahanta, Cristina, et al.
Veröffentlicht: (2025)
von: Mahanta, Cristina, et al.
Veröffentlicht: (2025)
Beyond Consensus: Perspectivist Modeling and Evaluation of Annotator Disagreement in NLP
von: Xu, Yinuo, et al.
Veröffentlicht: (2026)
von: Xu, Yinuo, et al.
Veröffentlicht: (2026)
Says Who? Effective Zero-Shot Annotation of Focalization
von: Hicke, Rebecca M. M., et al.
Veröffentlicht: (2024)
von: Hicke, Rebecca M. M., et al.
Veröffentlicht: (2024)
The Nature of NLP: Analyzing Contributions in NLP Papers
von: Pramanick, Aniket, et al.
Veröffentlicht: (2024)
von: Pramanick, Aniket, et al.
Veröffentlicht: (2024)
Blind Spots and Biases: Exploring the Role of Annotator Cognitive Biases in NLP
von: Gautam, Sanjana, et al.
Veröffentlicht: (2024)
von: Gautam, Sanjana, et al.
Veröffentlicht: (2024)
Argument Summarization and its Evaluation in the Era of Large Language Models
von: Altemeyer, Moritz, et al.
Veröffentlicht: (2025)
von: Altemeyer, Moritz, et al.
Veröffentlicht: (2025)
Disability-First AI Dataset Annotation: Co-designing Stuttered Speech Annotation Guidelines with People Who Stutter
von: Tang, Xinru, et al.
Veröffentlicht: (2026)
von: Tang, Xinru, et al.
Veröffentlicht: (2026)
Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
von: Eger, Steffen, et al.
Veröffentlicht: (2025)
von: Eger, Steffen, et al.
Veröffentlicht: (2025)
Counting on Consensus: Selecting the Right Inter-annotator Agreement Metric for NLP Annotation and Evaluation
von: James, Joseph
Veröffentlicht: (2026)
von: James, Joseph
Veröffentlicht: (2026)
Cross-lingual Cross-temporal Summarization: Dataset, Models, Evaluation
von: Zhang, Ran, et al.
Veröffentlicht: (2023)
von: Zhang, Ran, et al.
Veröffentlicht: (2023)
How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
von: Zhang, Ran, et al.
Veröffentlicht: (2024)
von: Zhang, Ran, et al.
Veröffentlicht: (2024)
AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ
von: Belouadi, Jonas, et al.
Veröffentlicht: (2023)
von: Belouadi, Jonas, et al.
Veröffentlicht: (2023)
Date Fragments: A Hidden Bottleneck of Tokenization for Temporal Reasoning
von: Bhatia, Gagan, et al.
Veröffentlicht: (2025)
von: Bhatia, Gagan, et al.
Veröffentlicht: (2025)
A Two-Year Perspective on Library Preservation: An Annotated Bibliography.
von: Fox, Lisa L.
Veröffentlicht: (1986)
von: Fox, Lisa L.
Veröffentlicht: (1986)
ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?
von: Zhang, Leixin, et al.
Veröffentlicht: (2024)
von: Zhang, Leixin, et al.
Veröffentlicht: (2024)
The Annotation Scarcity Paradox in Low-Resource NLP Evaluation: A Decade of Acceleration and Emerging Constraints
von: Marivate, Vukosi
Veröffentlicht: (2026)
von: Marivate, Vukosi
Veröffentlicht: (2026)
Ähnliche Einträge
-
Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics
von: Roy, Subhadeep, et al.
Veröffentlicht: (2026) -
TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning
von: Greisinger, Christian, et al.
Veröffentlicht: (2026) -
GerAV: Towards New Heights in German Authorship Verification using Fine-Tuned LLMs on a New Benchmark
von: Kiefer, Lotta, et al.
Veröffentlicht: (2026) -
DeepSeek-R1 vs. o3-mini: How Well can Reasoning LLMs Evaluate MT and Summarization?
von: Larionov, Daniil, et al.
Veröffentlicht: (2025) -
PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation
von: Leiter, Christoph, et al.
Veröffentlicht: (2024)