Counting on Consensus: Selecting the Right Inter-annotator Agreement Metric for NLP Annotation and Evaluation
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | James, Joseph |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Consensus: Perspectivist Modeling and Evaluation of Annotator Disagreement in NLP
von: Xu, Yinuo, et al.
Veröffentlicht: (2026)
von: Xu, Yinuo, et al.
Veröffentlicht: (2026)
Beyond Majority Voting: Agreement-Based Clustering to Model Annotator Perspectives in Subjective NLP Tasks
von: Belay, Tadesse Destaw, et al.
Veröffentlicht: (2026)
von: Belay, Tadesse Destaw, et al.
Veröffentlicht: (2026)
Evaluation Metrics for Text Data Augmentation in NLP
von: Amadeus, Marcellus, et al.
Veröffentlicht: (2024)
von: Amadeus, Marcellus, et al.
Veröffentlicht: (2024)
Exploring transfer learning for Deep NLP systems on rarely annotated languages
von: Yadav, Dipendra, et al.
Veröffentlicht: (2024)
von: Yadav, Dipendra, et al.
Veröffentlicht: (2024)
Estimating Agreement by Chance for Sequence Annotation
von: Li, Diya, et al.
Veröffentlicht: (2024)
von: Li, Diya, et al.
Veröffentlicht: (2024)
We Need to Talk About Classification Evaluation Metrics in NLP
von: Vickers, Peter, et al.
Veröffentlicht: (2024)
von: Vickers, Peter, et al.
Veröffentlicht: (2024)
LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation
von: Eigler, Lukáš, et al.
Veröffentlicht: (2026)
von: Eigler, Lukáš, et al.
Veröffentlicht: (2026)
A Decomposition-Based Approach for Evaluating and Analyzing Inter-Annotator Disagreement
von: Levi, Effi, et al.
Veröffentlicht: (2022)
von: Levi, Effi, et al.
Veröffentlicht: (2022)
Select, Label, Evaluate: Active Testing in NLP
von: Purificato, Antonio, et al.
Veröffentlicht: (2026)
von: Purificato, Antonio, et al.
Veröffentlicht: (2026)
Annotation alignment: Comparing LLM and human annotations of conversational safety
von: Movva, Rajiv, et al.
Veröffentlicht: (2024)
von: Movva, Rajiv, et al.
Veröffentlicht: (2024)
Annotator-Centric Active Learning for Subjective NLP Tasks
von: van der Meer, Michiel, et al.
Veröffentlicht: (2024)
von: van der Meer, Michiel, et al.
Veröffentlicht: (2024)
BiST: A Gold Standard Bangla-English Bilingual Corpus for Sentence Structure and Tense Classification with Inter-Annotator Agreement
von: Shafi, Abdullah Al, et al.
Veröffentlicht: (2026)
von: Shafi, Abdullah Al, et al.
Veröffentlicht: (2026)
The Annotation Scarcity Paradox in Low-Resource NLP Evaluation: A Decade of Acceleration and Emerging Constraints
von: Marivate, Vukosi
Veröffentlicht: (2026)
von: Marivate, Vukosi
Veröffentlicht: (2026)
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
Foundations and Evaluations in NLP
von: Park, Jungyeul
Veröffentlicht: (2025)
von: Park, Jungyeul
Veröffentlicht: (2025)
From Divergence to Consensus: Evaluating the Role of Large Language Models in Facilitating Agreement through Adaptive Strategies
von: Triantafyllopoulos, Loukas, et al.
Veröffentlicht: (2025)
von: Triantafyllopoulos, Loukas, et al.
Veröffentlicht: (2025)
K$α$LOS finds Consensus: A Meta-Algorithm for Evaluating Inter-Annotator Agreement in Complex Vision Tasks
von: Tschirschwitz, David, et al.
Veröffentlicht: (2026)
von: Tschirschwitz, David, et al.
Veröffentlicht: (2026)
"You are an expert annotator": Automatic Best-Worst-Scaling Annotations for Emotion Intensity Modeling
von: Bagdon, Christopher, et al.
Veröffentlicht: (2024)
von: Bagdon, Christopher, et al.
Veröffentlicht: (2024)
Blind Spots and Biases: Exploring the Role of Annotator Cognitive Biases in NLP
von: Gautam, Sanjana, et al.
Veröffentlicht: (2024)
von: Gautam, Sanjana, et al.
Veröffentlicht: (2024)
Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation
von: Sperber, Matthias, et al.
Veröffentlicht: (2024)
von: Sperber, Matthias, et al.
Veröffentlicht: (2024)
Consistency is Key: Disentangling Label Variation in Natural Language Processing with Intra-Annotator Agreement
von: Abercrombie, Gavin, et al.
Veröffentlicht: (2023)
von: Abercrombie, Gavin, et al.
Veröffentlicht: (2023)
Do LLMs Signal When They're Right? Evidence from Neuron Agreement
von: Chen, Kang, et al.
Veröffentlicht: (2025)
von: Chen, Kang, et al.
Veröffentlicht: (2025)
Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025
von: Kunilovskaya, Maria, et al.
Veröffentlicht: (2026)
von: Kunilovskaya, Maria, et al.
Veröffentlicht: (2026)
Structured Disagreement in Health-Literacy Annotation: Epistemic Stability, Conceptual Difficulty, and Agreement-Stratified Inference
von: Kellert, Olga, et al.
Veröffentlicht: (2026)
von: Kellert, Olga, et al.
Veröffentlicht: (2026)
CPE-Identifier: Automated CPE identification and CVE summaries annotation with Deep Learning and NLP
von: Hu, Wanyu, et al.
Veröffentlicht: (2024)
von: Hu, Wanyu, et al.
Veröffentlicht: (2024)
Beyond Weaponization: NLP Security for Medium and Lower-Resourced Languages in Their Own Right
von: Lent, Heather
Veröffentlicht: (2025)
von: Lent, Heather
Veröffentlicht: (2025)
LlamBERT: Large-scale low-cost data annotation in NLP
von: Csanády, Bálint, et al.
Veröffentlicht: (2024)
von: Csanády, Bálint, et al.
Veröffentlicht: (2024)
NLP for Local Governance Meeting Records: A Focus Article on Tasks, Datasets, Metrics and Benchmark
von: Campos, Ricardo, et al.
Veröffentlicht: (2026)
von: Campos, Ricardo, et al.
Veröffentlicht: (2026)
Metric-Dependent Annotation Saturation for Learning from Label Distributions
von: Kohli, Guneet
Veröffentlicht: (2026)
von: Kohli, Guneet
Veröffentlicht: (2026)
Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models
von: Amiri-Margavi, Alireza, et al.
Veröffentlicht: (2024)
von: Amiri-Margavi, Alireza, et al.
Veröffentlicht: (2024)
Privacy Evaluation Benchmarks for NLP Models
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
TurkicNLP: An NLP Toolkit for Turkic Languages
von: Hakimov, Sherzod
Veröffentlicht: (2026)
von: Hakimov, Sherzod
Veröffentlicht: (2026)
The Nature of NLP: Analyzing Contributions in NLP Papers
von: Pramanick, Aniket, et al.
Veröffentlicht: (2024)
von: Pramanick, Aniket, et al.
Veröffentlicht: (2024)
The Consensus Trap: Dissecting Subjectivity and the "Ground Truth" Illusion in Data Annotation
von: Munir, Sheza, et al.
Veröffentlicht: (2026)
von: Munir, Sheza, et al.
Veröffentlicht: (2026)
Beyond Catalogue Counts: the Dataset Visibility Asymmetry in Low-Resource Multilingual NLP
von: Tan, Zhiyin, et al.
Veröffentlicht: (2026)
von: Tan, Zhiyin, et al.
Veröffentlicht: (2026)
Not All Subjectivity Is the Same! Defining Desiderata for the Evaluation of Subjectivity in NLP
von: Khurana, Urja, et al.
Veröffentlicht: (2026)
von: Khurana, Urja, et al.
Veröffentlicht: (2026)
A Survey on Out-of-Distribution Evaluation of Neural NLP Models
von: Li, Xinzhe, et al.
Veröffentlicht: (2023)
von: Li, Xinzhe, et al.
Veröffentlicht: (2023)
AraReasoner: Evaluating Reasoning-Based LLMs for Arabic NLP
von: Hasanaath, Ahmed, et al.
Veröffentlicht: (2025)
von: Hasanaath, Ahmed, et al.
Veröffentlicht: (2025)
Disaggregation Reveals Hidden Training Dynamics: The Case of Agreement Attraction
von: Michaelov, James A., et al.
Veröffentlicht: (2025)
von: Michaelov, James A., et al.
Veröffentlicht: (2025)
Tokenization and Morphological Fidelity in Uralic NLP: A Cross-Lingual Evaluation
von: Xu, Nuo, et al.
Veröffentlicht: (2026)
von: Xu, Nuo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Beyond Consensus: Perspectivist Modeling and Evaluation of Annotator Disagreement in NLP
von: Xu, Yinuo, et al.
Veröffentlicht: (2026) -
Beyond Majority Voting: Agreement-Based Clustering to Model Annotator Perspectives in Subjective NLP Tasks
von: Belay, Tadesse Destaw, et al.
Veröffentlicht: (2026) -
Evaluation Metrics for Text Data Augmentation in NLP
von: Amadeus, Marcellus, et al.
Veröffentlicht: (2024) -
Exploring transfer learning for Deep NLP systems on rarely annotated languages
von: Yadav, Dipendra, et al.
Veröffentlicht: (2024) -
Estimating Agreement by Chance for Sequence Annotation
von: Li, Diya, et al.
Veröffentlicht: (2024)