Reconstructing Biological Pathways by Applying Selective Incremental Learning to (Very) Small Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Saha, Pranta, Reimer, Joyce, Byrns, Brook, Burbridge, Connor, Dhar, Neeraj, Chen, Jeffrey, Rayan, Steven, Broderick, Gordon
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918085121802240
author Saha, Pranta
Reimer, Joyce
Byrns, Brook
Burbridge, Connor
Dhar, Neeraj
Chen, Jeffrey
Rayan, Steven
Broderick, Gordon
author_facet Saha, Pranta
Reimer, Joyce
Byrns, Brook
Burbridge, Connor
Dhar, Neeraj
Chen, Jeffrey
Rayan, Steven
Broderick, Gordon
contents The use of generative artificial intelligence (AI) models is becoming ubiquitous in many fields. Though progress continues to be made, general purpose large language AI models (LLM) show a tendency to deliver creative answers, often called "hallucinations", which have slowed their application in the medical and biomedical fields where accuracy is paramount. We propose that the design and use of much smaller, domain and even task-specific LM may be a more rational and appropriate use of this technology in biomedical research. In this work we apply a very small LM by today's standards to the specialized task of predicting regulatory interactions between molecular components to fill gaps in our current understanding of intracellular pathways. Toward this we attempt to correctly posit known pathway-informed interactions recovered from manually curated pathway databases by selecting and using only the most informative examples as part of an active learning scheme. With this example we show that a small (~110 million parameters) LM based on a Bidirectional Encoder Representations from Transformers (BERT) architecture can propose molecular interactions relevant to tuberculosis persistence and transmission with over 80% accuracy using less than 25% of the ~520 regulatory relationships in question. Using information entropy as a metric for the iterative selection of new tuning examples, we also find that increased accuracy is driven by favoring the use of the incorrectly assigned statements with the highest certainty (lowest entropy). In contrast, the concurrent use of correct but least certain examples contributed little and may have even been detrimental to the learning rate.
format Preprint
id arxiv_https___arxiv_org_abs_2507_04432
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reconstructing Biological Pathways by Applying Selective Incremental Learning to (Very) Small Language Models
Saha, Pranta
Reimer, Joyce
Byrns, Brook
Burbridge, Connor
Dhar, Neeraj
Chen, Jeffrey
Rayan, Steven
Broderick, Gordon
Molecular Networks
Computation and Language
Information Theory
Machine Learning
Performance
The use of generative artificial intelligence (AI) models is becoming ubiquitous in many fields. Though progress continues to be made, general purpose large language AI models (LLM) show a tendency to deliver creative answers, often called "hallucinations", which have slowed their application in the medical and biomedical fields where accuracy is paramount. We propose that the design and use of much smaller, domain and even task-specific LM may be a more rational and appropriate use of this technology in biomedical research. In this work we apply a very small LM by today's standards to the specialized task of predicting regulatory interactions between molecular components to fill gaps in our current understanding of intracellular pathways. Toward this we attempt to correctly posit known pathway-informed interactions recovered from manually curated pathway databases by selecting and using only the most informative examples as part of an active learning scheme. With this example we show that a small (~110 million parameters) LM based on a Bidirectional Encoder Representations from Transformers (BERT) architecture can propose molecular interactions relevant to tuberculosis persistence and transmission with over 80% accuracy using less than 25% of the ~520 regulatory relationships in question. Using information entropy as a metric for the iterative selection of new tuning examples, we also find that increased accuracy is driven by favoring the use of the incorrectly assigned statements with the highest certainty (lowest entropy). In contrast, the concurrent use of correct but least certain examples contributed little and may have even been detrimental to the learning rate.
title Reconstructing Biological Pathways by Applying Selective Incremental Learning to (Very) Small Language Models
topic Molecular Networks
Computation and Language
Information Theory
Machine Learning
Performance
url https://arxiv.org/abs/2507.04432