Dimension-free Score Matching and Time Bootstrapping for Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumar, Syamantak, Nagaraj, Dheeraj, Sarkar, Purnamrita
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914112715358208
author Kumar, Syamantak
Nagaraj, Dheeraj
Sarkar, Purnamrita
author_facet Kumar, Syamantak
Nagaraj, Dheeraj
Sarkar, Purnamrita
contents Diffusion models generate samples by estimating the score function of the target distribution at various noise levels. The model is trained using samples drawn from the target distribution by progressively adding noise. Previous sample complexity bounds have polynomial dependence on the dimension $d$, apart from a $\log(|\mathcal{H}|)$ term, where $\mathcal{H}$ is the hypothesis class. In this work, we establish the first (nearly) dimension-free sample complexity bounds, modulo the $\log(|\mathcal{H}|)$ dependence, for learning these score functions, achieving a double exponential improvement in the dimension over prior results. A key aspect of our analysis is the use of a single function approximator to jointly estimate scores across noise levels, a practical feature that enables generalization across time steps. We introduce a martingale-based error decomposition and sharp variance bounds, enabling efficient learning from dependent data generated by Markov processes, which may be of independent interest. Building on these insights, we propose Bootstrapped Score Matching (BSM), a variance reduction technique that leverages previously learned scores to improve accuracy at higher noise levels. These results provide insights into the efficiency and effectiveness of diffusion models for generative modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2502_10354
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dimension-free Score Matching and Time Bootstrapping for Diffusion Models
Kumar, Syamantak
Nagaraj, Dheeraj
Sarkar, Purnamrita
Machine Learning
Statistics Theory
Diffusion models generate samples by estimating the score function of the target distribution at various noise levels. The model is trained using samples drawn from the target distribution by progressively adding noise. Previous sample complexity bounds have polynomial dependence on the dimension $d$, apart from a $\log(|\mathcal{H}|)$ term, where $\mathcal{H}$ is the hypothesis class. In this work, we establish the first (nearly) dimension-free sample complexity bounds, modulo the $\log(|\mathcal{H}|)$ dependence, for learning these score functions, achieving a double exponential improvement in the dimension over prior results. A key aspect of our analysis is the use of a single function approximator to jointly estimate scores across noise levels, a practical feature that enables generalization across time steps. We introduce a martingale-based error decomposition and sharp variance bounds, enabling efficient learning from dependent data generated by Markov processes, which may be of independent interest. Building on these insights, we propose Bootstrapped Score Matching (BSM), a variance reduction technique that leverages previously learned scores to improve accuracy at higher noise levels. These results provide insights into the efficiency and effectiveness of diffusion models for generative modeling.
title Dimension-free Score Matching and Time Bootstrapping for Diffusion Models
topic Machine Learning
Statistics Theory
url https://arxiv.org/abs/2502.10354