BioReason: Incentivizing Multimodal Biological Reasoning within a DNA-LLM Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fallahpour, Adibvafa, Magnuson, Andrew, Gupta, Purav, Ma, Shihao, Naimer, Jack, Shah, Arnav, Duan, Haonan, Ibrahim, Omar, Goodarzi, Hani, Maddison, Chris J., Wang, Bo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918169774391296
author Fallahpour, Adibvafa
Magnuson, Andrew
Gupta, Purav
Ma, Shihao
Naimer, Jack
Shah, Arnav
Duan, Haonan
Ibrahim, Omar
Goodarzi, Hani
Maddison, Chris J.
Wang, Bo
author_facet Fallahpour, Adibvafa
Magnuson, Andrew
Gupta, Purav
Ma, Shihao
Naimer, Jack
Shah, Arnav
Duan, Haonan
Ibrahim, Omar
Goodarzi, Hani
Maddison, Chris J.
Wang, Bo
contents Unlocking deep and interpretable biological reasoning from complex genomic data remains a major AI challenge limiting scientific progress. While current DNA foundation models excel at representing sequences, they struggle with multi-step reasoning and lack transparent, biologically meaningful explanations. BioReason addresses this by tightly integrating a DNA foundation model with a large language model (LLM), enabling the LLM to directly interpret and reason over genomic information. Through supervised fine-tuning and reinforcement learning, BioReason learns to produce logical, biologically coherent deductions. It achieves major performance gains, boosting KEGG-based disease pathway prediction accuracy from 86% to 98% and improving variant effect prediction by an average of 15% over strong baselines. BioReason can reason over unseen biological entities and explain its decisions step by step, offering a transformative framework for interpretable, mechanistic AI in biology. All data, code, and checkpoints are available at https://github.com/bowang-lab/BioReason
format Preprint
id arxiv_https___arxiv_org_abs_2505_23579
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BioReason: Incentivizing Multimodal Biological Reasoning within a DNA-LLM Model
Fallahpour, Adibvafa
Magnuson, Andrew
Gupta, Purav
Ma, Shihao
Naimer, Jack
Shah, Arnav
Duan, Haonan
Ibrahim, Omar
Goodarzi, Hani
Maddison, Chris J.
Wang, Bo
Machine Learning
J.3; I.2
Unlocking deep and interpretable biological reasoning from complex genomic data remains a major AI challenge limiting scientific progress. While current DNA foundation models excel at representing sequences, they struggle with multi-step reasoning and lack transparent, biologically meaningful explanations. BioReason addresses this by tightly integrating a DNA foundation model with a large language model (LLM), enabling the LLM to directly interpret and reason over genomic information. Through supervised fine-tuning and reinforcement learning, BioReason learns to produce logical, biologically coherent deductions. It achieves major performance gains, boosting KEGG-based disease pathway prediction accuracy from 86% to 98% and improving variant effect prediction by an average of 15% over strong baselines. BioReason can reason over unseen biological entities and explain its decisions step by step, offering a transformative framework for interpretable, mechanistic AI in biology. All data, code, and checkpoints are available at https://github.com/bowang-lab/BioReason
title BioReason: Incentivizing Multimodal Biological Reasoning within a DNA-LLM Model
topic Machine Learning
J.3; I.2
url https://arxiv.org/abs/2505.23579