Large Language Models with Human-In-The-Loop Validation for Systematic Review Data Extraction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Schroeder, Noah L., Jaldi, Chris Davis, Zhang, Shan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910792139407360
author Schroeder, Noah L.
Jaldi, Chris Davis
Zhang, Shan
author_facet Schroeder, Noah L.
Jaldi, Chris Davis
Zhang, Shan
contents Systematic reviews are time-consuming endeavors. Historically speaking, knowledgeable humans have had to screen and extract data from studies before it can be analyzed. However, large language models (LLMs) hold promise to greatly accelerate this process. After a pilot study which showed great promise, we investigated the use of freely available LLMs for extracting data for systematic reviews. Using three different LLMs, we extracted 24 types of data, 9 explicitly stated variables and 15 derived categorical variables, from 112 studies that were included in a published scoping review. Overall we found that Gemini 1.5 Flash, Gemini 1.5 Pro, and Mistral Large 2 performed reasonably well, with 71.17%, 72.14%, and 62.43% of data extracted being consistent with human coding, respectively. While promising, these results highlight the dire need for a human-in-the-loop (HIL) process for AI-assisted data extraction. As a result, we present a free, open-source program we developed (AIDE) to facilitate user-friendly, HIL data extraction with LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2501_11840
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large Language Models with Human-In-The-Loop Validation for Systematic Review Data Extraction
Schroeder, Noah L.
Jaldi, Chris Davis
Zhang, Shan
Human-Computer Interaction
Systematic reviews are time-consuming endeavors. Historically speaking, knowledgeable humans have had to screen and extract data from studies before it can be analyzed. However, large language models (LLMs) hold promise to greatly accelerate this process. After a pilot study which showed great promise, we investigated the use of freely available LLMs for extracting data for systematic reviews. Using three different LLMs, we extracted 24 types of data, 9 explicitly stated variables and 15 derived categorical variables, from 112 studies that were included in a published scoping review. Overall we found that Gemini 1.5 Flash, Gemini 1.5 Pro, and Mistral Large 2 performed reasonably well, with 71.17%, 72.14%, and 62.43% of data extracted being consistent with human coding, respectively. While promising, these results highlight the dire need for a human-in-the-loop (HIL) process for AI-assisted data extraction. As a result, we present a free, open-source program we developed (AIDE) to facilitate user-friendly, HIL data extraction with LLMs.
title Large Language Models with Human-In-The-Loop Validation for Systematic Review Data Extraction
topic Human-Computer Interaction
url https://arxiv.org/abs/2501.11840