BioLangFusion: Multimodal Fusion of DNA, mRNA, and Protein Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mollaysa, Amina, Moskale, Artem, Pati, Pushpak, Mansi, Tommaso, Prakash, Mangal, Liao, Rui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908402201919488
author Mollaysa, Amina
Moskale, Artem
Pati, Pushpak
Mansi, Tommaso
Prakash, Mangal
Liao, Rui
author_facet Mollaysa, Amina
Moskale, Artem
Pati, Pushpak
Mansi, Tommaso
Prakash, Mangal
Liao, Rui
contents We present BioLangFusion, a simple approach for integrating pre-trained DNA, mRNA, and protein language models into unified molecular representations. Motivated by the central dogma of molecular biology (information flow from gene to transcript to protein), we align per-modality embeddings at the biologically meaningful codon level (three nucleotides encoding one amino acid) to ensure direct cross-modal correspondence. BioLangFusion studies three standard fusion techniques: (i) codon-level embedding concatenation, (ii) entropy-regularized attention pooling inspired by multiple-instance learning, and (iii) cross-modal multi-head attention -- each technique providing a different inductive bias for combining modality-specific signals. These methods require no additional pre-training or modification of the base models, allowing straightforward integration with existing sequence-based foundation models. Across five molecular property prediction tasks, BioLangFusion outperforms strong unimodal baselines, showing that even simple fusion of pre-trained models can capture complementary multi-omic information with minimal overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08936
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BioLangFusion: Multimodal Fusion of DNA, mRNA, and Protein Language Models
Mollaysa, Amina
Moskale, Artem
Pati, Pushpak
Mansi, Tommaso
Prakash, Mangal
Liao, Rui
Machine Learning
We present BioLangFusion, a simple approach for integrating pre-trained DNA, mRNA, and protein language models into unified molecular representations. Motivated by the central dogma of molecular biology (information flow from gene to transcript to protein), we align per-modality embeddings at the biologically meaningful codon level (three nucleotides encoding one amino acid) to ensure direct cross-modal correspondence. BioLangFusion studies three standard fusion techniques: (i) codon-level embedding concatenation, (ii) entropy-regularized attention pooling inspired by multiple-instance learning, and (iii) cross-modal multi-head attention -- each technique providing a different inductive bias for combining modality-specific signals. These methods require no additional pre-training or modification of the base models, allowing straightforward integration with existing sequence-based foundation models. Across five molecular property prediction tasks, BioLangFusion outperforms strong unimodal baselines, showing that even simple fusion of pre-trained models can capture complementary multi-omic information with minimal overhead.
title BioLangFusion: Multimodal Fusion of DNA, mRNA, and Protein Language Models
topic Machine Learning
url https://arxiv.org/abs/2506.08936