Development and Benchmarking of a Blended Human-AI Qualitative Research Assistant

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Matveyenko, Joseph, Liu, James, Parsons, John David, Brown, Ryan A., Palimaru, Alina, Puri, Prateek
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915674332332032
author Matveyenko, Joseph
Liu, James
Parsons, John David
Brown, Ryan A.
Palimaru, Alina
Puri, Prateek
author_facet Matveyenko, Joseph
Liu, James
Parsons, John David
Brown, Ryan A.
Palimaru, Alina
Puri, Prateek
contents Qualitative research emphasizes constructing meaning through iterative engagement with textual data. Traditionally this human-driven process requires navigating coder fatigue and interpretative drift, thus posing challenges when scaling analysis to larger, more complex datasets. Computational approaches to augment qualitative research have been met with skepticism, partly due to their inability to replicate the nuance, context-awareness, and sophistication of human analysis. Large language models, however, present new opportunities to automate aspects of qualitative analysis while upholding rigor and research quality in important ways. To assess their benefits and limitations - and build trust among qualitative researchers - these approaches must be rigorously benchmarked against human-generated datasets. In this work, we benchmark Muse, an interactive, AI-powered qualitative research system that allows researchers to identify themes and annotate datasets, finding an inter-rater reliability between Muse and humans of Cohen's $κ$ = 0.71 for well-specified codes. We also conduct robust error analysis to identify failure mode, guide future improvements, and demonstrate the capacity to correct for human bias.
format Preprint
id arxiv_https___arxiv_org_abs_2512_00009
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Development and Benchmarking of a Blended Human-AI Qualitative Research Assistant
Matveyenko, Joseph
Liu, James
Parsons, John David
Brown, Ryan A.
Palimaru, Alina
Puri, Prateek
Human-Computer Interaction
Artificial Intelligence
I.2.7
Qualitative research emphasizes constructing meaning through iterative engagement with textual data. Traditionally this human-driven process requires navigating coder fatigue and interpretative drift, thus posing challenges when scaling analysis to larger, more complex datasets. Computational approaches to augment qualitative research have been met with skepticism, partly due to their inability to replicate the nuance, context-awareness, and sophistication of human analysis. Large language models, however, present new opportunities to automate aspects of qualitative analysis while upholding rigor and research quality in important ways. To assess their benefits and limitations - and build trust among qualitative researchers - these approaches must be rigorously benchmarked against human-generated datasets. In this work, we benchmark Muse, an interactive, AI-powered qualitative research system that allows researchers to identify themes and annotate datasets, finding an inter-rater reliability between Muse and humans of Cohen's $κ$ = 0.71 for well-specified codes. We also conduct robust error analysis to identify failure mode, guide future improvements, and demonstrate the capacity to correct for human bias.
title Development and Benchmarking of a Blended Human-AI Qualitative Research Assistant
topic Human-Computer Interaction
Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2512.00009