Elucidating the Design Space of Multimodal Protein Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hsieh, Cheng-Yen, Wang, Xinyou, Zhang, Daiheng, Xue, Dongyu, Ye, Fei, Huang, Shujian, Zheng, Zaixiang, Gu, Quanquan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911001141575680
author Hsieh, Cheng-Yen
Wang, Xinyou
Zhang, Daiheng
Xue, Dongyu
Ye, Fei
Huang, Shujian
Zheng, Zaixiang
Gu, Quanquan
author_facet Hsieh, Cheng-Yen
Wang, Xinyou
Zhang, Daiheng
Xue, Dongyu
Ye, Fei
Huang, Shujian
Zheng, Zaixiang
Gu, Quanquan
contents Multimodal protein language models (PLMs) integrate sequence and token-based structural information, serving as a powerful foundation for protein modeling, generation, and design. However, the reliance on tokenizing 3D structures into discrete tokens causes substantial loss of fidelity about fine-grained structural details and correlations. In this paper, we systematically elucidate the design space of multimodal PLMs to overcome their limitations. We identify tokenization loss and inaccurate structure token predictions by the PLMs as major bottlenecks. To address these, our proposed design space covers improved generative modeling, structure-aware architectures and representation learning, and data exploration. Our advancements approach finer-grained supervision, demonstrating that token-based multimodal PLMs can achieve robust structural modeling. The effective design methods dramatically improve the structure generation diversity, and notably, folding abilities of our 650M model by reducing the RMSD from 5.52 to 2.36 on PDB testset, even outperforming 3B baselines and on par with the specialized folding models. Project page and code: https://bytedance.github.io/dplm/dplm-2.1/.
format Preprint
id arxiv_https___arxiv_org_abs_2504_11454
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Elucidating the Design Space of Multimodal Protein Language Models
Hsieh, Cheng-Yen
Wang, Xinyou
Zhang, Daiheng
Xue, Dongyu
Ye, Fei
Huang, Shujian
Zheng, Zaixiang
Gu, Quanquan
Machine Learning
Artificial Intelligence
Quantitative Methods
Multimodal protein language models (PLMs) integrate sequence and token-based structural information, serving as a powerful foundation for protein modeling, generation, and design. However, the reliance on tokenizing 3D structures into discrete tokens causes substantial loss of fidelity about fine-grained structural details and correlations. In this paper, we systematically elucidate the design space of multimodal PLMs to overcome their limitations. We identify tokenization loss and inaccurate structure token predictions by the PLMs as major bottlenecks. To address these, our proposed design space covers improved generative modeling, structure-aware architectures and representation learning, and data exploration. Our advancements approach finer-grained supervision, demonstrating that token-based multimodal PLMs can achieve robust structural modeling. The effective design methods dramatically improve the structure generation diversity, and notably, folding abilities of our 650M model by reducing the RMSD from 5.52 to 2.36 on PDB testset, even outperforming 3B baselines and on par with the specialized folding models. Project page and code: https://bytedance.github.io/dplm/dplm-2.1/.
title Elucidating the Design Space of Multimodal Protein Language Models
topic Machine Learning
Artificial Intelligence
Quantitative Methods
url https://arxiv.org/abs/2504.11454