VibeProteinBench: An Evaluation Benchmark for Language-interfaced Vibe Protein Design

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Seo, Hyunjin, Ahn, Hongjoon, Park, Jimin, Han, Sungjun, Lee, Gyubok, Yang, Soojung, Brown, Joseph S, Chen, Leo, Nesr, Gina El, Eweje, Feyisayo, Gurev, Sarah, Lee, Hyejin, Liu, Cheng-Hao, Liu, Junlang, Qi, Zhihui, Lee, Gyu Rie, Ahn, Sungsoo, Shin, Jamin, Jung, Sangwon
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917504770637824
author Seo, Hyunjin
Ahn, Hongjoon
Park, Jimin
Han, Sungjun
Lee, Gyubok
Yang, Soojung
Brown, Joseph S
Chen, Leo
Nesr, Gina El
Eweje, Feyisayo
Gurev, Sarah
Lee, Hyejin
Liu, Cheng-Hao
Liu, Junlang
Qi, Zhihui
Lee, Gyu Rie
Ahn, Sungsoo
Shin, Jamin
Jung, Sangwon
author_facet Seo, Hyunjin
Ahn, Hongjoon
Park, Jimin
Han, Sungjun
Lee, Gyubok
Yang, Soojung
Brown, Joseph S
Chen, Leo
Nesr, Gina El
Eweje, Feyisayo
Gurev, Sarah
Lee, Hyejin
Liu, Cheng-Hao
Liu, Junlang
Qi, Zhihui
Lee, Gyu Rie
Ahn, Sungsoo
Shin, Jamin
Jung, Sangwon
contents Protein design aims to compose amino-acid sequences that fold into stable three-dimensional structures while satisfying targeted functional properties. The field is increasingly shifting toward vibe protein design, where a single model is expected to generate novel sequences, engineer existing proteins, and reason about protein characteristics through flexible natural-language constraints. Large language models (LLMs) have emerged as a leading paradigm in this space. However, existing evaluation benchmarks often limit their scope to a partial aspect of protein design, while others restrict design objectives to structured input schemas, lacking an integrated framework that evaluates the broad spectrum of protein design competence under open-ended intents. To this end, we present Vibe Protein design Benchmark (VibeProteinBench), a language-interfaced benchmark that probes generalist capabilities through three complementary stages mirroring a computational protein design workflow: recognition, engineering, and generation. Each stage is grounded in expert-curated mechanistic rationales and multi-faceted in silico validation, to computationally verify whether model outputs are biologically plausible. Evaluations across diverse general-purpose and domain-specialized LLMs reveal that no model achieves strong performance across all three stages, suggesting that generalist protein design remains a substantial open challenge for current LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10978
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VibeProteinBench: An Evaluation Benchmark for Language-interfaced Vibe Protein Design
Seo, Hyunjin
Ahn, Hongjoon
Park, Jimin
Han, Sungjun
Lee, Gyubok
Yang, Soojung
Brown, Joseph S
Chen, Leo
Nesr, Gina El
Eweje, Feyisayo
Gurev, Sarah
Lee, Hyejin
Liu, Cheng-Hao
Liu, Junlang
Qi, Zhihui
Lee, Gyu Rie
Ahn, Sungsoo
Shin, Jamin
Jung, Sangwon
Quantitative Methods
Protein design aims to compose amino-acid sequences that fold into stable three-dimensional structures while satisfying targeted functional properties. The field is increasingly shifting toward vibe protein design, where a single model is expected to generate novel sequences, engineer existing proteins, and reason about protein characteristics through flexible natural-language constraints. Large language models (LLMs) have emerged as a leading paradigm in this space. However, existing evaluation benchmarks often limit their scope to a partial aspect of protein design, while others restrict design objectives to structured input schemas, lacking an integrated framework that evaluates the broad spectrum of protein design competence under open-ended intents. To this end, we present Vibe Protein design Benchmark (VibeProteinBench), a language-interfaced benchmark that probes generalist capabilities through three complementary stages mirroring a computational protein design workflow: recognition, engineering, and generation. Each stage is grounded in expert-curated mechanistic rationales and multi-faceted in silico validation, to computationally verify whether model outputs are biologically plausible. Evaluations across diverse general-purpose and domain-specialized LLMs reveal that no model achieves strong performance across all three stages, suggesting that generalist protein design remains a substantial open challenge for current LLMs.
title VibeProteinBench: An Evaluation Benchmark for Language-interfaced Vibe Protein Design
topic Quantitative Methods
url https://arxiv.org/abs/2605.10978