To Err Is Human; To Annotate, SILICON? Toward Robust Reproducibility in LLM Annotation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cheng, Xiang, Mayya, Raveesh, Sedoc, João |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VariErr NLI: Separating Annotation Error from Human Label Variation
von: Weber-Genzel, Leon, et al.
Veröffentlicht: (2024)
von: Weber-Genzel, Leon, et al.
Veröffentlicht: (2024)
MedErrBench: A Fine-Grained Multilingual Benchmark for Medical Error Detection and Correction with Clinical Expert Annotations
von: Ma, Congbo, et al.
Veröffentlicht: (2026)
von: Ma, Congbo, et al.
Veröffentlicht: (2026)
Hybrid Annotation for Propaganda Detection: Integrating LLM Pre-Annotations with Human Intelligence
von: Sahitaj, Ariana, et al.
Veröffentlicht: (2025)
von: Sahitaj, Ariana, et al.
Veröffentlicht: (2025)
To Err Is Human, but Llamas Can Learn It Too
von: Luhtaru, Agnes, et al.
Veröffentlicht: (2024)
von: Luhtaru, Agnes, et al.
Veröffentlicht: (2024)
Refining and Reusing Annotation Guidelines for LLM Annotation
von: Kim, Kon Woo, et al.
Veröffentlicht: (2026)
von: Kim, Kon Woo, et al.
Veröffentlicht: (2026)
MEGAnno+: A Human-LLM Collaborative Annotation System
von: Kim, Hannah, et al.
Veröffentlicht: (2024)
von: Kim, Hannah, et al.
Veröffentlicht: (2024)
To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis
von: Bianchi, Federico, et al.
Veröffentlicht: (2025)
von: Bianchi, Federico, et al.
Veröffentlicht: (2025)
DBOT: Artificial Intelligence for Systematic Long-Term Investing
von: Dhar, Vasant, et al.
Veröffentlicht: (2025)
von: Dhar, Vasant, et al.
Veröffentlicht: (2025)
Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and Targets
von: Giorgi, Tommaso, et al.
Veröffentlicht: (2024)
von: Giorgi, Tommaso, et al.
Veröffentlicht: (2024)
LLM-as-an-Annotator: Training Lightweight Models with LLM-Annotated Examples for Aspect Sentiment Tuple Prediction
von: Hellwig, Nils Constantin, et al.
Veröffentlicht: (2026)
von: Hellwig, Nils Constantin, et al.
Veröffentlicht: (2026)
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
von: Calderon, Nitay, et al.
Veröffentlicht: (2025)
von: Calderon, Nitay, et al.
Veröffentlicht: (2025)
It's Difficult to be Neutral -- Human and LLM-based Sentiment Annotation of Patient Comments
von: Mæhlum, Petter, et al.
Veröffentlicht: (2024)
von: Mæhlum, Petter, et al.
Veröffentlicht: (2024)
Human-LLM Hybrid Text Answer Aggregation for Crowd Annotations
von: Li, Jiyi
Veröffentlicht: (2024)
von: Li, Jiyi
Veröffentlicht: (2024)
Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
von: Kim, Kon Woo, et al.
Veröffentlicht: (2025)
von: Kim, Kon Woo, et al.
Veröffentlicht: (2025)
Who speaks like a style of Vitamin: Towards Syntax-Aware DialogueSummarization using Multi-task Learning
von: Lee, Seolhwa, et al.
Veröffentlicht: (2021)
von: Lee, Seolhwa, et al.
Veröffentlicht: (2021)
Towards Generating Automatic Anaphora Annotations
von: Taji, Dima, et al.
Veröffentlicht: (2025)
von: Taji, Dima, et al.
Veröffentlicht: (2025)
Reasoning and the Trusting Behavior of DeepSeek and GPT: An Experiment Revealing Hidden Fault Lines in Large Language Models
von: Li, Rubing, et al.
Veröffentlicht: (2025)
von: Li, Rubing, et al.
Veröffentlicht: (2025)
Towards Event Extraction with Massive Types: LLM-based Collaborative Annotation and Partitioning Extraction
von: Liu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Liu, Wenxuan, et al.
Veröffentlicht: (2025)
ReasonScaffold: A Scaffolded Reasoning-based Annotation Protocol for Human-AI Co-Annotation
von: Sudheendra, Smitha Muthya, et al.
Veröffentlicht: (2026)
von: Sudheendra, Smitha Muthya, et al.
Veröffentlicht: (2026)
Large Human Language Models: A Need and the Challenges
von: Soni, Nikita, et al.
Veröffentlicht: (2023)
von: Soni, Nikita, et al.
Veröffentlicht: (2023)
AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM Annotators
von: Ni, Jingwei, et al.
Veröffentlicht: (2024)
von: Ni, Jingwei, et al.
Veröffentlicht: (2024)
Large Language Models Are Effective Human Annotation Assistants, But Not Good Independent Annotators
von: Gu, Feng, et al.
Veröffentlicht: (2025)
von: Gu, Feng, et al.
Veröffentlicht: (2025)
Large Language Models Reproduce Racial Stereotypes When Used for Text Annotation
von: Törnberg, Petter
Veröffentlicht: (2026)
von: Törnberg, Petter
Veröffentlicht: (2026)
Human-Annotated NER Dataset for the Kyrgyz Language
von: Turatali, Timur, et al.
Veröffentlicht: (2025)
von: Turatali, Timur, et al.
Veröffentlicht: (2025)
An Annotated Reading of 'The Singer of Tales' in the LLM Era
von: Varshney, Kush R.
Veröffentlicht: (2025)
von: Varshney, Kush R.
Veröffentlicht: (2025)
CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data Annotation
von: Li, Minzhi, et al.
Veröffentlicht: (2023)
von: Li, Minzhi, et al.
Veröffentlicht: (2023)
Prompt-Counterfactual Explanations for Generative AI System Behavior
von: Goethals, Sofie, et al.
Veröffentlicht: (2026)
von: Goethals, Sofie, et al.
Veröffentlicht: (2026)
GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction
von: Felkner, Virginia K., et al.
Veröffentlicht: (2024)
von: Felkner, Virginia K., et al.
Veröffentlicht: (2024)
Towards Automating Text Annotation: A Case Study on Semantic Proximity Annotation using GPT-4
von: Yadav, Sachin, et al.
Veröffentlicht: (2024)
von: Yadav, Sachin, et al.
Veröffentlicht: (2024)
WhoSaidIt: Human-LLM Collaborative Annotation for Text-Based Multilingual Speaker-Attribute Classification
von: Gao, Lingyu, et al.
Veröffentlicht: (2026)
von: Gao, Lingyu, et al.
Veröffentlicht: (2026)
LATA: A Tool for LLM-Assisted Translation Annotation
von: Huang, Baorong, et al.
Veröffentlicht: (2026)
von: Huang, Baorong, et al.
Veröffentlicht: (2026)
Evaluating the Impact of LLM-Assisted Annotation in a Perspectivized Setting: the Case of FrameNet Annotation
von: Belcavello, Frederico, et al.
Veröffentlicht: (2025)
von: Belcavello, Frederico, et al.
Veröffentlicht: (2025)
Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI
von: Pangakis, Nicholas, et al.
Veröffentlicht: (2024)
von: Pangakis, Nicholas, et al.
Veröffentlicht: (2024)
Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization
von: Chen, Beiduo, et al.
Veröffentlicht: (2026)
von: Chen, Beiduo, et al.
Veröffentlicht: (2026)
Baby Bear: Seeking a Just Right Rating Scale for Scalar Annotations
von: Han, Xu, et al.
Veröffentlicht: (2024)
von: Han, Xu, et al.
Veröffentlicht: (2024)
From Human Annotation to Automation: LLM-in-the-Loop Active Learning for Arabic Sentiment Analysis
von: Refai, Dania, et al.
Veröffentlicht: (2025)
von: Refai, Dania, et al.
Veröffentlicht: (2025)
On Limitations of LLM as Annotator for Low Resource Languages
von: Jadhav, Suramya, et al.
Veröffentlicht: (2024)
von: Jadhav, Suramya, et al.
Veröffentlicht: (2024)
Emergent Convergence in Multi-Agent LLM Annotation
von: Parfenova, Angelina, et al.
Veröffentlicht: (2025)
von: Parfenova, Angelina, et al.
Veröffentlicht: (2025)
Toward Generalized Cross-Lingual Hateful Language Detection with Web-Scale Data and Ensemble LLM Annotations
von: Dang, Dang H., et al.
Veröffentlicht: (2026)
von: Dang, Dang H., et al.
Veröffentlicht: (2026)
Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025
von: Kunilovskaya, Maria, et al.
Veröffentlicht: (2026)
von: Kunilovskaya, Maria, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
VariErr NLI: Separating Annotation Error from Human Label Variation
von: Weber-Genzel, Leon, et al.
Veröffentlicht: (2024) -
MedErrBench: A Fine-Grained Multilingual Benchmark for Medical Error Detection and Correction with Clinical Expert Annotations
von: Ma, Congbo, et al.
Veröffentlicht: (2026) -
Hybrid Annotation for Propaganda Detection: Integrating LLM Pre-Annotations with Human Intelligence
von: Sahitaj, Ariana, et al.
Veröffentlicht: (2025) -
To Err Is Human, but Llamas Can Learn It Too
von: Luhtaru, Agnes, et al.
Veröffentlicht: (2024) -
Refining and Reusing Annotation Guidelines for LLM Annotation
von: Kim, Kon Woo, et al.
Veröffentlicht: (2026)