FoldSAE: Learning to Steer Protein Folding Through Sparse Representations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zarzecki, Wojciech, Szymczak, Paulina, Szczurek, Ewa, Deja, Kamil
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914173461463040
author Zarzecki, Wojciech
Szymczak, Paulina
Szczurek, Ewa
Deja, Kamil
author_facet Zarzecki, Wojciech
Szymczak, Paulina
Szczurek, Ewa
Deja, Kamil
contents RFdiffusion is a popular and well-established model for generation of protein structures. However, this generative process offers limited insight into its internal representations and how they contribute to the final protein structure. Concurrently, recent work in mechanistic interpretability has successfully used Sparse Autoencoders (SAEs) to discover interpretable features within neural networks. We combine these concepts by applying SAE to the internal representations of RFdiffusion to uncover secondary structure-specific features and establish a relationship between them and generated protein structures. Building on these insights, we introduce a novel steering mechanism that enables precise control of secondary structure formation through a tunable hyperparameter, while simultaneously revealing interpretable block and neuron-level representations within RFdiffusion. Our work pioneers a new framework for making RFdiffusion more interpretable, demonstrating how understanding internal features can be directly translated into precise control over the protein design process.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22519
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FoldSAE: Learning to Steer Protein Folding Through Sparse Representations
Zarzecki, Wojciech
Szymczak, Paulina
Szczurek, Ewa
Deja, Kamil
Quantitative Methods
RFdiffusion is a popular and well-established model for generation of protein structures. However, this generative process offers limited insight into its internal representations and how they contribute to the final protein structure. Concurrently, recent work in mechanistic interpretability has successfully used Sparse Autoencoders (SAEs) to discover interpretable features within neural networks. We combine these concepts by applying SAE to the internal representations of RFdiffusion to uncover secondary structure-specific features and establish a relationship between them and generated protein structures. Building on these insights, we introduce a novel steering mechanism that enables precise control of secondary structure formation through a tunable hyperparameter, while simultaneously revealing interpretable block and neuron-level representations within RFdiffusion. Our work pioneers a new framework for making RFdiffusion more interpretable, demonstrating how understanding internal features can be directly translated into precise control over the protein design process.
title FoldSAE: Learning to Steer Protein Folding Through Sparse Representations
topic Quantitative Methods
url https://arxiv.org/abs/2511.22519