Salvato in:
Dettagli Bibliografici
Autori principali: Chhetri, Tek Raj, Chen, Yibei, Trivedi, Puja, Jarecka, Dorota, Haobsh, Saif, Ray, Patrick, Ng, Lydia, Ghosh, Satrajit S.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:https://arxiv.org/abs/2507.03674
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911702756360192
author Chhetri, Tek Raj
Chen, Yibei
Trivedi, Puja
Jarecka, Dorota
Haobsh, Saif
Ray, Patrick
Ng, Lydia
Ghosh, Satrajit S.
author_facet Chhetri, Tek Raj
Chen, Yibei
Trivedi, Puja
Jarecka, Dorota
Haobsh, Saif
Ray, Patrick
Ng, Lydia
Ghosh, Satrajit S.
contents Extracting structured information from scientific literature is critical for accelerating discovery, yet Large Language Models (LLMs) often struggle in specialized domains that require expert knowledge and generalize poorly across tasks. We introduce \textsc{StructSense}, a modular, task-agnostic, open-source framework that integrates ontology-guided symbolic knowledge, agentic self-evaluative refinement, and human-in-the-loop validation for robust domain-aware extraction. We evaluate \textsc{StructSense} on three tasks of increasing semantic complexity: schema-based extraction of assessment instruments (91--100\% accuracy), metadata and resource extraction from scientific papers (86--93\% overall), and named entity recognition (NER) from neuroscience literature (58--75\% label accuracy across 8,882 entities). On two biomedical NER benchmarks (NCBI Disease and S800 Species), the system achieves $\geq$90\% relaxed recall and 62.5--85.8\% strict recall while extracting 1,000--3,600 additional entities beyond gold annotations. The local concept mapping service achieves Hits@1 of 62--82\% under strict matching and 68--86\% under semantic matching. These results across three domains demonstrate that \textsc{StructSense} generalizes across tasks while maintaining source grounding and provenance transparency.
format Preprint
id arxiv_https___arxiv_org_abs_2507_03674
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle STRUCTSENSE: A Task-Agnostic Agentic Framework for Structured Information Extraction with Human-In-The-Loop Evaluation and Benchmarking
Chhetri, Tek Raj
Chen, Yibei
Trivedi, Puja
Jarecka, Dorota
Haobsh, Saif
Ray, Patrick
Ng, Lydia
Ghosh, Satrajit S.
Computation and Language
Artificial Intelligence
Extracting structured information from scientific literature is critical for accelerating discovery, yet Large Language Models (LLMs) often struggle in specialized domains that require expert knowledge and generalize poorly across tasks. We introduce \textsc{StructSense}, a modular, task-agnostic, open-source framework that integrates ontology-guided symbolic knowledge, agentic self-evaluative refinement, and human-in-the-loop validation for robust domain-aware extraction. We evaluate \textsc{StructSense} on three tasks of increasing semantic complexity: schema-based extraction of assessment instruments (91--100\% accuracy), metadata and resource extraction from scientific papers (86--93\% overall), and named entity recognition (NER) from neuroscience literature (58--75\% label accuracy across 8,882 entities). On two biomedical NER benchmarks (NCBI Disease and S800 Species), the system achieves $\geq$90\% relaxed recall and 62.5--85.8\% strict recall while extracting 1,000--3,600 additional entities beyond gold annotations. The local concept mapping service achieves Hits@1 of 62--82\% under strict matching and 68--86\% under semantic matching. These results across three domains demonstrate that \textsc{StructSense} generalizes across tasks while maintaining source grounding and provenance transparency.
title STRUCTSENSE: A Task-Agnostic Agentic Framework for Structured Information Extraction with Human-In-The-Loop Evaluation and Benchmarking
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2507.03674