Text Annotation via Inductive Coding: Comparing Human Experts to LLMs in Qualitative Data Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Parfenova, Angelina, Marfurt, Andreas, Denzler, Alexander, Pfeffer, Juergen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914174437687296
author Parfenova, Angelina
Marfurt, Andreas
Denzler, Alexander
Pfeffer, Juergen
author_facet Parfenova, Angelina
Marfurt, Andreas
Denzler, Alexander
Pfeffer, Juergen
contents This paper investigates the automation of qualitative data analysis, focusing on inductive coding using large language models (LLMs). Unlike traditional approaches that rely on deductive methods with predefined labels, this research investigates the inductive process where labels emerge from the data. The study evaluates the performance of six open-source LLMs compared to human experts. As part of the evaluation, experts rated the perceived difficulty of the quotes they coded. The results reveal a peculiar dichotomy: human coders consistently perform well when labeling complex sentences but struggle with simpler ones, while LLMs exhibit the opposite trend. Additionally, the study explores systematic deviations in both human and LLM generated labels by comparing them to the golden standard from the test set. While human annotations may sometimes differ from the golden standard, they are often rated more favorably by other humans. In contrast, some LLMs demonstrate closer alignment with the true labels but receive lower evaluations from experts.
format Preprint
id arxiv_https___arxiv_org_abs_2512_00046
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Text Annotation via Inductive Coding: Comparing Human Experts to LLMs in Qualitative Data Analysis
Parfenova, Angelina
Marfurt, Andreas
Denzler, Alexander
Pfeffer, Juergen
Computation and Language
Artificial Intelligence
This paper investigates the automation of qualitative data analysis, focusing on inductive coding using large language models (LLMs). Unlike traditional approaches that rely on deductive methods with predefined labels, this research investigates the inductive process where labels emerge from the data. The study evaluates the performance of six open-source LLMs compared to human experts. As part of the evaluation, experts rated the perceived difficulty of the quotes they coded. The results reveal a peculiar dichotomy: human coders consistently perform well when labeling complex sentences but struggle with simpler ones, while LLMs exhibit the opposite trend. Additionally, the study explores systematic deviations in both human and LLM generated labels by comparing them to the golden standard from the test set. While human annotations may sometimes differ from the golden standard, they are often rated more favorably by other humans. In contrast, some LLMs demonstrate closer alignment with the true labels but receive lower evaluations from experts.
title Text Annotation via Inductive Coding: Comparing Human Experts to LLMs in Qualitative Data Analysis
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.00046