Protein Structure Tokenization: Benchmarking and New Recipe

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yuan, Xinyu, Wang, Zichen, Collins, Marcus, Rangwala, Huzefa
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913910966190080
author Yuan, Xinyu
Wang, Zichen
Collins, Marcus
Rangwala, Huzefa
author_facet Yuan, Xinyu
Wang, Zichen
Collins, Marcus
Rangwala, Huzefa
contents Recent years have witnessed a surge in the development of protein structural tokenization methods, which chunk protein 3D structures into discrete or continuous representations. Structure tokenization enables the direct application of powerful techniques like language modeling for protein structures, and large multimodal models to integrate structures with protein sequences and functional texts. Despite the progress, the capabilities and limitations of these methods remain poorly understood due to the lack of a unified evaluation framework. We first introduce StructTokenBench, a framework that comprehensively evaluates the quality and efficiency of structure tokenizers, focusing on fine-grained local substructures rather than global structures, as typical in existing benchmarks. Our evaluations reveal that no single model dominates all benchmarking perspectives. Observations of codebook under-utilization led us to develop AminoAseed, a simple yet effective strategy that enhances codebook gradient updates and optimally balances codebook size and dimension for improved tokenizer utilization and quality. Compared to the leading model ESM3, our method achieves an average of 6.31% performance improvement across 24 supervised tasks, with sensitivity and utilization rates increased by 12.83% and 124.03%, respectively. Source code and model weights are available at https://github.com/KatarinaYuan/StructTokenBench
format Preprint
id arxiv_https___arxiv_org_abs_2503_00089
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Protein Structure Tokenization: Benchmarking and New Recipe
Yuan, Xinyu
Wang, Zichen
Collins, Marcus
Rangwala, Huzefa
Quantitative Methods
Artificial Intelligence
Machine Learning
Recent years have witnessed a surge in the development of protein structural tokenization methods, which chunk protein 3D structures into discrete or continuous representations. Structure tokenization enables the direct application of powerful techniques like language modeling for protein structures, and large multimodal models to integrate structures with protein sequences and functional texts. Despite the progress, the capabilities and limitations of these methods remain poorly understood due to the lack of a unified evaluation framework. We first introduce StructTokenBench, a framework that comprehensively evaluates the quality and efficiency of structure tokenizers, focusing on fine-grained local substructures rather than global structures, as typical in existing benchmarks. Our evaluations reveal that no single model dominates all benchmarking perspectives. Observations of codebook under-utilization led us to develop AminoAseed, a simple yet effective strategy that enhances codebook gradient updates and optimally balances codebook size and dimension for improved tokenizer utilization and quality. Compared to the leading model ESM3, our method achieves an average of 6.31% performance improvement across 24 supervised tasks, with sensitivity and utilization rates increased by 12.83% and 124.03%, respectively. Source code and model weights are available at https://github.com/KatarinaYuan/StructTokenBench
title Protein Structure Tokenization: Benchmarking and New Recipe
topic Quantitative Methods
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2503.00089