BioLangFusion: Multimodal Fusion of DNA, mRNA, and Protein Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908402201919488 |
|---|---|
| author | Mollaysa, Amina Moskale, Artem Pati, Pushpak Mansi, Tommaso Prakash, Mangal Liao, Rui |
| author_facet | Mollaysa, Amina Moskale, Artem Pati, Pushpak Mansi, Tommaso Prakash, Mangal Liao, Rui |
| contents | We present BioLangFusion, a simple approach for integrating pre-trained DNA, mRNA, and protein language models into unified molecular representations. Motivated by the central dogma of molecular biology (information flow from gene to transcript to protein), we align per-modality embeddings at the biologically meaningful codon level (three nucleotides encoding one amino acid) to ensure direct cross-modal correspondence. BioLangFusion studies three standard fusion techniques: (i) codon-level embedding concatenation, (ii) entropy-regularized attention pooling inspired by multiple-instance learning, and (iii) cross-modal multi-head attention -- each technique providing a different inductive bias for combining modality-specific signals. These methods require no additional pre-training or modification of the base models, allowing straightforward integration with existing sequence-based foundation models. Across five molecular property prediction tasks, BioLangFusion outperforms strong unimodal baselines, showing that even simple fusion of pre-trained models can capture complementary multi-omic information with minimal overhead. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_08936 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | BioLangFusion: Multimodal Fusion of DNA, mRNA, and Protein Language Models Mollaysa, Amina Moskale, Artem Pati, Pushpak Mansi, Tommaso Prakash, Mangal Liao, Rui Machine Learning We present BioLangFusion, a simple approach for integrating pre-trained DNA, mRNA, and protein language models into unified molecular representations. Motivated by the central dogma of molecular biology (information flow from gene to transcript to protein), we align per-modality embeddings at the biologically meaningful codon level (three nucleotides encoding one amino acid) to ensure direct cross-modal correspondence. BioLangFusion studies three standard fusion techniques: (i) codon-level embedding concatenation, (ii) entropy-regularized attention pooling inspired by multiple-instance learning, and (iii) cross-modal multi-head attention -- each technique providing a different inductive bias for combining modality-specific signals. These methods require no additional pre-training or modification of the base models, allowing straightforward integration with existing sequence-based foundation models. Across five molecular property prediction tasks, BioLangFusion outperforms strong unimodal baselines, showing that even simple fusion of pre-trained models can capture complementary multi-omic information with minimal overhead. |
| title | BioLangFusion: Multimodal Fusion of DNA, mRNA, and Protein Language Models |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2506.08936 |