Protein FID: Improved Evaluation of Protein Structure Generative Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Faltings, Felix, Stark, Hannes, Jaakkola, Tommi, Barzilay, Regina
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916860256059392
author Faltings, Felix
Stark, Hannes
Jaakkola, Tommi
Barzilay, Regina
author_facet Faltings, Felix
Stark, Hannes
Jaakkola, Tommi
Barzilay, Regina
contents Protein structure generative models have seen a recent surge of interest, but meaningfully evaluating them computationally is an active area of research. While current metrics have driven useful progress, they do not capture how well models sample the design space represented by the training data. We argue for a protein Frechet Inception Distance (FID) metric to supplement current evaluations with a measure of distributional similarity in a semantically meaningful latent space. Our FID behaves desirably under protein structure perturbations and correctly recapitulates similarities between protein samples: it correlates with optimal transport distances and recovers FoldSeek clusters and the CATH hierarchy. Evaluating current protein structure generative models with FID shows that they fall short of modeling the distribution of PDB proteins.
format Preprint
id arxiv_https___arxiv_org_abs_2505_08041
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Protein FID: Improved Evaluation of Protein Structure Generative Models
Faltings, Felix
Stark, Hannes
Jaakkola, Tommi
Barzilay, Regina
Biomolecules
Protein structure generative models have seen a recent surge of interest, but meaningfully evaluating them computationally is an active area of research. While current metrics have driven useful progress, they do not capture how well models sample the design space represented by the training data. We argue for a protein Frechet Inception Distance (FID) metric to supplement current evaluations with a measure of distributional similarity in a semantically meaningful latent space. Our FID behaves desirably under protein structure perturbations and correctly recapitulates similarities between protein samples: it correlates with optimal transport distances and recovers FoldSeek clusters and the CATH hierarchy. Evaluating current protein structure generative models with FID shows that they fall short of modeling the distribution of PDB proteins.
title Protein FID: Improved Evaluation of Protein Structure Generative Models
topic Biomolecules
url https://arxiv.org/abs/2505.08041