SeedFold: Scaling Biomolecular Structure Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Yi, Lu, Chan, Ma, Yiming, Qu, Wei, Ye, Fei, Zhang, Kexin, Wang, Lan, Gui, Minrui, Gu, Quanquan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908741854560256
author Zhou, Yi
Lu, Chan
Ma, Yiming
Qu, Wei
Ye, Fei
Zhang, Kexin
Wang, Lan
Gui, Minrui
Gu, Quanquan
author_facet Zhou, Yi
Lu, Chan
Ma, Yiming
Qu, Wei
Ye, Fei
Zhang, Kexin
Wang, Lan
Gui, Minrui
Gu, Quanquan
contents Highly accurate biomolecular structure prediction is a key component of developing biomolecular foundation models, and one of the most critical aspects of building foundation models is identifying the recipes for scaling the model. In this work, we present SeedFold, a folding model that successfully scales up the model capacity. Our contributions are threefold: first, we identify an effective width-scaling strategy for the Pairformer to increase representation capacity; second, we introduce a novel linear triangular attention that reduces computational complexity to enable efficient scaling; finally, we construct a large-scale distillation dataset to substantially enlarge the training set. Experiments on FoldBench show that SeedFold outperforms AlphaFold3 on most protein-related tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2512_24354
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SeedFold: Scaling Biomolecular Structure Prediction
Zhou, Yi
Lu, Chan
Ma, Yiming
Qu, Wei
Ye, Fei
Zhang, Kexin
Wang, Lan
Gui, Minrui
Gu, Quanquan
Biomolecules
Highly accurate biomolecular structure prediction is a key component of developing biomolecular foundation models, and one of the most critical aspects of building foundation models is identifying the recipes for scaling the model. In this work, we present SeedFold, a folding model that successfully scales up the model capacity. Our contributions are threefold: first, we identify an effective width-scaling strategy for the Pairformer to increase representation capacity; second, we introduce a novel linear triangular attention that reduces computational complexity to enable efficient scaling; finally, we construct a large-scale distillation dataset to substantially enlarge the training set. Experiments on FoldBench show that SeedFold outperforms AlphaFold3 on most protein-related tasks.
title SeedFold: Scaling Biomolecular Structure Prediction
topic Biomolecules
url https://arxiv.org/abs/2512.24354