LLMs as Data Annotators: How Close Are We to Human Performance
Fuente:
arXiv
Saved in:
| Main Authors: | Haq, Muhammad Uzair Ul, Rigoni, Davide, Sperduti, Alessandro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Text to Talent: A Pipeline for Extracting Insights from Candidate Profiles
by: Frazzetto, Paolo, et al.
Published: (2025)
by: Frazzetto, Paolo, et al.
Published: (2025)
APT-Pipe: A Prompt-Tuning Tool for Social Data Annotation using ChatGPT
by: Zhu, Yiming, et al.
Published: (2024)
by: Zhu, Yiming, et al.
Published: (2024)
A Neural Rewriting System to Solve Algorithmic Problems
by: Petruzzellis, Flavio, et al.
Published: (2024)
by: Petruzzellis, Flavio, et al.
Published: (2024)
Benchmarking GPT-4 on Algorithmic Problems: A Systematic Evaluation of Prompting Strategies
by: Petruzzellis, Flavio, et al.
Published: (2024)
by: Petruzzellis, Flavio, et al.
Published: (2024)
Typoglycemia under the Hood: Investigating Language Models' Understanding of Scrambled Words
by: Sperduti, Gianluca, et al.
Published: (2025)
by: Sperduti, Gianluca, et al.
Published: (2025)
Assessing the Emergent Symbolic Reasoning Abilities of Llama Large Language Models
by: Petruzzellis, Flavio, et al.
Published: (2024)
by: Petruzzellis, Flavio, et al.
Published: (2024)
Misspellings in Natural Language Processing: A survey
by: Sperduti, Gianluca, et al.
Published: (2025)
by: Sperduti, Gianluca, et al.
Published: (2025)
How Far Are We? Systematic Evaluation of LLMs vs. Human Experts in Mathematical Contest in Modeling
by: Liu, Yuhang, et al.
Published: (2026)
by: Liu, Yuhang, et al.
Published: (2026)
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
by: Calderon, Nitay, et al.
Published: (2025)
by: Calderon, Nitay, et al.
Published: (2025)
LLMs as Span Annotators: A Comparative Study of LLMs and Humans
by: Kasner, Zdeněk, et al.
Published: (2025)
by: Kasner, Zdeněk, et al.
Published: (2025)
Promptception: How Sensitive Are Large Multimodal Models to Prompts?
by: Ismithdeen, Mohamed Insaf, et al.
Published: (2025)
by: Ismithdeen, Mohamed Insaf, et al.
Published: (2025)
LLMs for Relational Reasoning: How Far are We?
by: Li, Zhiming, et al.
Published: (2024)
by: Li, Zhiming, et al.
Published: (2024)
Audio-Based Crowd-Sourced Evaluation of Machine Translation Quality
by: Haq, Sami Ul, et al.
Published: (2025)
by: Haq, Sami Ul, et al.
Published: (2025)
Exploring the Capability of ChatGPT to Reproduce Human Labels for Social Computing Tasks (Extended Version)
by: Zhu, Yiming, et al.
Published: (2024)
by: Zhu, Yiming, et al.
Published: (2024)
How Far Are We From AGI: Are LLMs All We Need?
by: Feng, Tao, et al.
Published: (2024)
by: Feng, Tao, et al.
Published: (2024)
Model Editing for LLMs4Code: How Far are We?
by: Li, Xiaopeng, et al.
Published: (2024)
by: Li, Xiaopeng, et al.
Published: (2024)
Text Annotation via Inductive Coding: Comparing Human Experts to LLMs in Qualitative Data Analysis
by: Parfenova, Angelina, et al.
Published: (2025)
by: Parfenova, Angelina, et al.
Published: (2025)
HIKMA: Human-Inspired Knowledge by Machine Agents through a Multi-Agent Framework for Semi-Autonomous Scientific Conferences
by: Tariq, Zain Ul Abideen, et al.
Published: (2025)
by: Tariq, Zain Ul Abideen, et al.
Published: (2025)
Personas with Attitudes: Controlling LLMs for Diverse Data Annotation
by: Fröhling, Leon, et al.
Published: (2024)
by: Fröhling, Leon, et al.
Published: (2024)
Fearful Falcons and Angry Llamas: Emotion Category Annotations of Arguments by Humans and LLMs
by: Greschner, Lynn, et al.
Published: (2024)
by: Greschner, Lynn, et al.
Published: (2024)
Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection
by: Hakimi, Ahmad Dawar, et al.
Published: (2026)
by: Hakimi, Ahmad Dawar, et al.
Published: (2026)
Position: AI Evaluation Should Learn from How We Test Humans
by: Zhuang, Yan, et al.
Published: (2023)
by: Zhuang, Yan, et al.
Published: (2023)
Claim Check-Worthiness Detection: How Well do LLMs Grasp Annotation Guidelines?
by: Majer, Laura, et al.
Published: (2024)
by: Majer, Laura, et al.
Published: (2024)
Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains
by: Wu, Juncheng, et al.
Published: (2025)
by: Wu, Juncheng, et al.
Published: (2025)
Where Are We? Evaluating LLM Performance on African Languages
by: Adebara, Ife, et al.
Published: (2025)
by: Adebara, Ife, et al.
Published: (2025)
LLM Compression: How Far Can We Go in Balancing Size and Performance?
by: Sk, Sahil, et al.
Published: (2025)
by: Sk, Sahil, et al.
Published: (2025)
Improving the OOD Performance of Closed-Source LLMs on NLI Through Strategic Data Selection
by: Stacey, Joe, et al.
Published: (2025)
by: Stacey, Joe, et al.
Published: (2025)
CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data Annotation
by: Li, Minzhi, et al.
Published: (2023)
by: Li, Minzhi, et al.
Published: (2023)
How Much of Your Data Can Suck? Thresholds for Domain Performance and Emergent Misalignment in LLMs
by: Ouyang, Jian, et al.
Published: (2025)
by: Ouyang, Jian, et al.
Published: (2025)
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
by: Huang, Jen-tse, et al.
Published: (2024)
by: Huang, Jen-tse, et al.
Published: (2024)
Evaluating Webcam-based Gaze Data as an Alternative for Human Rationale Annotations
by: Brandl, Stephanie, et al.
Published: (2024)
by: Brandl, Stephanie, et al.
Published: (2024)
Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
by: Atil, Berk, et al.
Published: (2025)
by: Atil, Berk, et al.
Published: (2025)
Performance Gains of LLMs With Humans in a World of LLMs Versus Humans
by: McCullum, Lucas, et al.
Published: (2025)
by: McCullum, Lucas, et al.
Published: (2025)
From Fallback to Frontline: When Can LLMs be Superior Annotators of Human Perspectives?
by: Amin, Hasan, et al.
Published: (2026)
by: Amin, Hasan, et al.
Published: (2026)
SynthDetoxM: Modern LLMs are Few-Shot Parallel Detoxification Data Annotators
by: Moskovskiy, Daniil, et al.
Published: (2025)
by: Moskovskiy, Daniil, et al.
Published: (2025)
Using LLMs to Aid Annotation and Collection of Clinically-Enriched Data in Bipolar Disorder and Schizophrenia
by: Aich, Ankit, et al.
Published: (2024)
by: Aich, Ankit, et al.
Published: (2024)
Long-context Reference-based MT Quality Estimation
by: Haq, Sami Ul, et al.
Published: (2025)
by: Haq, Sami Ul, et al.
Published: (2025)
When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction
by: Dongre, Vardhan, et al.
Published: (2026)
by: Dongre, Vardhan, et al.
Published: (2026)
Should We Respect LLMs? A Cross-Lingual Study on the Influence of Prompt Politeness on LLM Performance
by: Yin, Ziqi, et al.
Published: (2024)
by: Yin, Ziqi, et al.
Published: (2024)
To Err Is Human; To Annotate, SILICON? Toward Robust Reproducibility in LLM Annotation
by: Cheng, Xiang, et al.
Published: (2024)
by: Cheng, Xiang, et al.
Published: (2024)
Similar Items
-
From Text to Talent: A Pipeline for Extracting Insights from Candidate Profiles
by: Frazzetto, Paolo, et al.
Published: (2025) -
APT-Pipe: A Prompt-Tuning Tool for Social Data Annotation using ChatGPT
by: Zhu, Yiming, et al.
Published: (2024) -
A Neural Rewriting System to Solve Algorithmic Problems
by: Petruzzellis, Flavio, et al.
Published: (2024) -
Benchmarking GPT-4 on Algorithmic Problems: A Systematic Evaluation of Prompting Strategies
by: Petruzzellis, Flavio, et al.
Published: (2024) -
Typoglycemia under the Hood: Investigating Language Models' Understanding of Scrambled Words
by: Sperduti, Gianluca, et al.
Published: (2025)