Data-Dependent Smoothing for Protein Discovery with Walk-Jump Sampling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Anumasa, Srinivas, C, Barath Chandran., Chen, Tingting, Liu, Dianbo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916930010480640
author Anumasa, Srinivas
C, Barath Chandran.
Chen, Tingting
Liu, Dianbo
author_facet Anumasa, Srinivas
C, Barath Chandran.
Chen, Tingting
Liu, Dianbo
contents Diffusion models have emerged as a powerful class of generative models by learning to iteratively reverse the noising process. Their ability to generate high-quality samples has extended beyond high-dimensional image data to other complex domains such as proteins, where data distributions are typically sparse and unevenly spread. Importantly, the sparsity itself is uneven. Empirically, we observed that while a small fraction of samples lie in dense clusters, the majority occupy regions of varying sparsity across the data space. Existing approaches largely ignore this data-dependent variability. In this work, we introduce a Data-Dependent Smoothing Walk-Jump framework that employs kernel density estimation (KDE) as a preprocessing step to estimate the noise scale $σ$ for each data point, followed by training a score model with these data-dependent $σ$ values. By incorporating local data geometry into the denoising process, our method accounts for the heterogeneous distribution of protein data. Empirical evaluations demonstrate that our approach yields consistent improvements across multiple metrics, highlighting the importance of data-aware sigma prediction for generative modeling in sparse, high-dimensional settings.
format Preprint
id arxiv_https___arxiv_org_abs_2509_02069
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Data-Dependent Smoothing for Protein Discovery with Walk-Jump Sampling
Anumasa, Srinivas
C, Barath Chandran.
Chen, Tingting
Liu, Dianbo
Machine Learning
Quantitative Methods
Diffusion models have emerged as a powerful class of generative models by learning to iteratively reverse the noising process. Their ability to generate high-quality samples has extended beyond high-dimensional image data to other complex domains such as proteins, where data distributions are typically sparse and unevenly spread. Importantly, the sparsity itself is uneven. Empirically, we observed that while a small fraction of samples lie in dense clusters, the majority occupy regions of varying sparsity across the data space. Existing approaches largely ignore this data-dependent variability. In this work, we introduce a Data-Dependent Smoothing Walk-Jump framework that employs kernel density estimation (KDE) as a preprocessing step to estimate the noise scale $σ$ for each data point, followed by training a score model with these data-dependent $σ$ values. By incorporating local data geometry into the denoising process, our method accounts for the heterogeneous distribution of protein data. Empirical evaluations demonstrate that our approach yields consistent improvements across multiple metrics, highlighting the importance of data-aware sigma prediction for generative modeling in sparse, high-dimensional settings.
title Data-Dependent Smoothing for Protein Discovery with Walk-Jump Sampling
topic Machine Learning
Quantitative Methods
url https://arxiv.org/abs/2509.02069