To Err Is Human; To Annotate, SILICON? Toward Robust Reproducibility in LLM Annotation
Fuente:
arXiv
Saved in:
| Main Authors: | Cheng, Xiang, Mayya, Raveesh, Sedoc, João |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VariErr NLI: Separating Annotation Error from Human Label Variation
by: Weber-Genzel, Leon, et al.
Published: (2024)
by: Weber-Genzel, Leon, et al.
Published: (2024)
MedErrBench: A Fine-Grained Multilingual Benchmark for Medical Error Detection and Correction with Clinical Expert Annotations
by: Ma, Congbo, et al.
Published: (2026)
by: Ma, Congbo, et al.
Published: (2026)
Hybrid Annotation for Propaganda Detection: Integrating LLM Pre-Annotations with Human Intelligence
by: Sahitaj, Ariana, et al.
Published: (2025)
by: Sahitaj, Ariana, et al.
Published: (2025)
To Err Is Human, but Llamas Can Learn It Too
by: Luhtaru, Agnes, et al.
Published: (2024)
by: Luhtaru, Agnes, et al.
Published: (2024)
Refining and Reusing Annotation Guidelines for LLM Annotation
by: Kim, Kon Woo, et al.
Published: (2026)
by: Kim, Kon Woo, et al.
Published: (2026)
MEGAnno+: A Human-LLM Collaborative Annotation System
by: Kim, Hannah, et al.
Published: (2024)
by: Kim, Hannah, et al.
Published: (2024)
To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis
by: Bianchi, Federico, et al.
Published: (2025)
by: Bianchi, Federico, et al.
Published: (2025)
DBOT: Artificial Intelligence for Systematic Long-Term Investing
by: Dhar, Vasant, et al.
Published: (2025)
by: Dhar, Vasant, et al.
Published: (2025)
Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and Targets
by: Giorgi, Tommaso, et al.
Published: (2024)
by: Giorgi, Tommaso, et al.
Published: (2024)
LLM-as-an-Annotator: Training Lightweight Models with LLM-Annotated Examples for Aspect Sentiment Tuple Prediction
by: Hellwig, Nils Constantin, et al.
Published: (2026)
by: Hellwig, Nils Constantin, et al.
Published: (2026)
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
by: Calderon, Nitay, et al.
Published: (2025)
by: Calderon, Nitay, et al.
Published: (2025)
It's Difficult to be Neutral -- Human and LLM-based Sentiment Annotation of Patient Comments
by: Mæhlum, Petter, et al.
Published: (2024)
by: Mæhlum, Petter, et al.
Published: (2024)
Human-LLM Hybrid Text Answer Aggregation for Crowd Annotations
by: Li, Jiyi
Published: (2024)
by: Li, Jiyi
Published: (2024)
Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
by: Kim, Kon Woo, et al.
Published: (2025)
by: Kim, Kon Woo, et al.
Published: (2025)
Who speaks like a style of Vitamin: Towards Syntax-Aware DialogueSummarization using Multi-task Learning
by: Lee, Seolhwa, et al.
Published: (2021)
by: Lee, Seolhwa, et al.
Published: (2021)
Towards Generating Automatic Anaphora Annotations
by: Taji, Dima, et al.
Published: (2025)
by: Taji, Dima, et al.
Published: (2025)
Reasoning and the Trusting Behavior of DeepSeek and GPT: An Experiment Revealing Hidden Fault Lines in Large Language Models
by: Li, Rubing, et al.
Published: (2025)
by: Li, Rubing, et al.
Published: (2025)
Towards Event Extraction with Massive Types: LLM-based Collaborative Annotation and Partitioning Extraction
by: Liu, Wenxuan, et al.
Published: (2025)
by: Liu, Wenxuan, et al.
Published: (2025)
ReasonScaffold: A Scaffolded Reasoning-based Annotation Protocol for Human-AI Co-Annotation
by: Sudheendra, Smitha Muthya, et al.
Published: (2026)
by: Sudheendra, Smitha Muthya, et al.
Published: (2026)
Large Human Language Models: A Need and the Challenges
by: Soni, Nikita, et al.
Published: (2023)
by: Soni, Nikita, et al.
Published: (2023)
AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM Annotators
by: Ni, Jingwei, et al.
Published: (2024)
by: Ni, Jingwei, et al.
Published: (2024)
Large Language Models Are Effective Human Annotation Assistants, But Not Good Independent Annotators
by: Gu, Feng, et al.
Published: (2025)
by: Gu, Feng, et al.
Published: (2025)
Large Language Models Reproduce Racial Stereotypes When Used for Text Annotation
by: Törnberg, Petter
Published: (2026)
by: Törnberg, Petter
Published: (2026)
Human-Annotated NER Dataset for the Kyrgyz Language
by: Turatali, Timur, et al.
Published: (2025)
by: Turatali, Timur, et al.
Published: (2025)
An Annotated Reading of 'The Singer of Tales' in the LLM Era
by: Varshney, Kush R.
Published: (2025)
by: Varshney, Kush R.
Published: (2025)
CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data Annotation
by: Li, Minzhi, et al.
Published: (2023)
by: Li, Minzhi, et al.
Published: (2023)
Prompt-Counterfactual Explanations for Generative AI System Behavior
by: Goethals, Sofie, et al.
Published: (2026)
by: Goethals, Sofie, et al.
Published: (2026)
GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction
by: Felkner, Virginia K., et al.
Published: (2024)
by: Felkner, Virginia K., et al.
Published: (2024)
Towards Automating Text Annotation: A Case Study on Semantic Proximity Annotation using GPT-4
by: Yadav, Sachin, et al.
Published: (2024)
by: Yadav, Sachin, et al.
Published: (2024)
WhoSaidIt: Human-LLM Collaborative Annotation for Text-Based Multilingual Speaker-Attribute Classification
by: Gao, Lingyu, et al.
Published: (2026)
by: Gao, Lingyu, et al.
Published: (2026)
LATA: A Tool for LLM-Assisted Translation Annotation
by: Huang, Baorong, et al.
Published: (2026)
by: Huang, Baorong, et al.
Published: (2026)
Evaluating the Impact of LLM-Assisted Annotation in a Perspectivized Setting: the Case of FrameNet Annotation
by: Belcavello, Frederico, et al.
Published: (2025)
by: Belcavello, Frederico, et al.
Published: (2025)
Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI
by: Pangakis, Nicholas, et al.
Published: (2024)
by: Pangakis, Nicholas, et al.
Published: (2024)
Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization
by: Chen, Beiduo, et al.
Published: (2026)
by: Chen, Beiduo, et al.
Published: (2026)
Baby Bear: Seeking a Just Right Rating Scale for Scalar Annotations
by: Han, Xu, et al.
Published: (2024)
by: Han, Xu, et al.
Published: (2024)
From Human Annotation to Automation: LLM-in-the-Loop Active Learning for Arabic Sentiment Analysis
by: Refai, Dania, et al.
Published: (2025)
by: Refai, Dania, et al.
Published: (2025)
On Limitations of LLM as Annotator for Low Resource Languages
by: Jadhav, Suramya, et al.
Published: (2024)
by: Jadhav, Suramya, et al.
Published: (2024)
Emergent Convergence in Multi-Agent LLM Annotation
by: Parfenova, Angelina, et al.
Published: (2025)
by: Parfenova, Angelina, et al.
Published: (2025)
Toward Generalized Cross-Lingual Hateful Language Detection with Web-Scale Data and Ensemble LLM Annotations
by: Dang, Dang H., et al.
Published: (2026)
by: Dang, Dang H., et al.
Published: (2026)
Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025
by: Kunilovskaya, Maria, et al.
Published: (2026)
by: Kunilovskaya, Maria, et al.
Published: (2026)
Similar Items
-
VariErr NLI: Separating Annotation Error from Human Label Variation
by: Weber-Genzel, Leon, et al.
Published: (2024) -
MedErrBench: A Fine-Grained Multilingual Benchmark for Medical Error Detection and Correction with Clinical Expert Annotations
by: Ma, Congbo, et al.
Published: (2026) -
Hybrid Annotation for Propaganda Detection: Integrating LLM Pre-Annotations with Human Intelligence
by: Sahitaj, Ariana, et al.
Published: (2025) -
To Err Is Human, but Llamas Can Learn It Too
by: Luhtaru, Agnes, et al.
Published: (2024) -
Refining and Reusing Annotation Guidelines for LLM Annotation
by: Kim, Kon Woo, et al.
Published: (2026)