SciDaSynth: Interactive Structured Data Extraction from Scientific Literature with Large Language Model

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Xingbo, Huey, Samantha L., Sheng, Rui, Mehta, Saurabh, Wang, Fei
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915599037235200
author Wang, Xingbo
Huey, Samantha L.
Sheng, Rui
Mehta, Saurabh
Wang, Fei
author_facet Wang, Xingbo
Huey, Samantha L.
Sheng, Rui
Mehta, Saurabh
Wang, Fei
contents The explosion of scientific literature has made the efficient and accurate extraction of structured data a critical component for advancing scientific knowledge and supporting evidence-based decision-making. However, existing tools often struggle to extract and structure multimodal, varied, and inconsistent information across documents into standardized formats. We introduce SciDaSynth, a novel interactive system powered by large language models (LLMs) that automatically generates structured data tables according to users' queries by integrating information from diverse sources, including text, tables, and figures. Furthermore, SciDaSynth supports efficient table data validation and refinement, featuring multi-faceted visual summaries and semantic grouping capabilities to resolve cross-document data inconsistencies. A within-subjects study with nutrition and NLP researchers demonstrates SciDaSynth's effectiveness in producing high-quality structured data more efficiently than baseline methods. We discuss design implications for human-AI collaborative systems supporting data extraction tasks. The system code is available at https://github.com/xingbow/SciDaEx
format Preprint
id arxiv_https___arxiv_org_abs_2404_13765
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SciDaSynth: Interactive Structured Data Extraction from Scientific Literature with Large Language Model
Wang, Xingbo
Huey, Samantha L.
Sheng, Rui
Mehta, Saurabh
Wang, Fei
Human-Computer Interaction
Computation and Language
The explosion of scientific literature has made the efficient and accurate extraction of structured data a critical component for advancing scientific knowledge and supporting evidence-based decision-making. However, existing tools often struggle to extract and structure multimodal, varied, and inconsistent information across documents into standardized formats. We introduce SciDaSynth, a novel interactive system powered by large language models (LLMs) that automatically generates structured data tables according to users' queries by integrating information from diverse sources, including text, tables, and figures. Furthermore, SciDaSynth supports efficient table data validation and refinement, featuring multi-faceted visual summaries and semantic grouping capabilities to resolve cross-document data inconsistencies. A within-subjects study with nutrition and NLP researchers demonstrates SciDaSynth's effectiveness in producing high-quality structured data more efficiently than baseline methods. We discuss design implications for human-AI collaborative systems supporting data extraction tasks. The system code is available at https://github.com/xingbow/SciDaEx
title SciDaSynth: Interactive Structured Data Extraction from Scientific Literature with Large Language Model
topic Human-Computer Interaction
Computation and Language
url https://arxiv.org/abs/2404.13765