Extracting Information in a Low-resource Setting: Case Study on Bioinformatics Workflows

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sebe, Clémence, Cohen-Boulakia, Sarah, Ferret, Olivier, Névéol, Aurélie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917949465427968
author Sebe, Clémence
Cohen-Boulakia, Sarah
Ferret, Olivier
Névéol, Aurélie
author_facet Sebe, Clémence
Cohen-Boulakia, Sarah
Ferret, Olivier
Névéol, Aurélie
contents Bioinformatics workflows are essential for complex biological data analyses and are often described in scientific articles with source code in public repositories. Extracting detailed workflow information from articles can improve accessibility and reusability but is hindered by limited annotated corpora. To address this, we framed the problem as a low-resource extraction task and tested four strategies: 1) creating a tailored annotated corpus, 2) few-shot named-entity recognition (NER) with an autoregressive language model, 3) NER using masked language models with existing and new corpora, and 4) integrating workflow knowledge into NER models. Using BioToFlow, a new corpus of 52 articles annotated with 16 entities, a SciBERT-based NER model achieved a 70.4 F-measure, comparable to inter-annotator agreement. While knowledge integration improved performance for specific entities, it was less effective across the entire information schema. Our results demonstrate that high-performance information extraction for bioinformatics workflows is achievable.
format Preprint
id arxiv_https___arxiv_org_abs_2411_19295
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Extracting Information in a Low-resource Setting: Case Study on Bioinformatics Workflows
Sebe, Clémence
Cohen-Boulakia, Sarah
Ferret, Olivier
Névéol, Aurélie
Computation and Language
Bioinformatics workflows are essential for complex biological data analyses and are often described in scientific articles with source code in public repositories. Extracting detailed workflow information from articles can improve accessibility and reusability but is hindered by limited annotated corpora. To address this, we framed the problem as a low-resource extraction task and tested four strategies: 1) creating a tailored annotated corpus, 2) few-shot named-entity recognition (NER) with an autoregressive language model, 3) NER using masked language models with existing and new corpora, and 4) integrating workflow knowledge into NER models. Using BioToFlow, a new corpus of 52 articles annotated with 16 entities, a SciBERT-based NER model achieved a 70.4 F-measure, comparable to inter-annotator agreement. While knowledge integration improved performance for specific entities, it was less effective across the entire information schema. Our results demonstrate that high-performance information extraction for bioinformatics workflows is achievable.
title Extracting Information in a Low-resource Setting: Case Study on Bioinformatics Workflows
topic Computation and Language
url https://arxiv.org/abs/2411.19295