Transformers Learn the Optimal DDPM Denoiser for Multi-Token GMMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Hongkang, Min, Hancheng, Vidal, Rene |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Understanding Incremental Learning with Closed-form Solution to Gradient Flow on Overparamerterized Matrix Factorization
von: Min, Hancheng, et al.
Veröffentlicht: (2025)
von: Min, Hancheng, et al.
Veröffentlicht: (2025)
Can Implicit Bias Imply Adversarial Robustness?
von: Min, Hancheng, et al.
Veröffentlicht: (2024)
von: Min, Hancheng, et al.
Veröffentlicht: (2024)
Early Neuron Alignment in Two-layer ReLU Networks with Small Initialization
von: Min, Hancheng, et al.
Veröffentlicht: (2023)
von: Min, Hancheng, et al.
Veröffentlicht: (2023)
Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data
von: Min, Hancheng, et al.
Veröffentlicht: (2025)
von: Min, Hancheng, et al.
Veröffentlicht: (2025)
Optimal Convergence Analysis of DDPM for General Distributions
von: Jiao, Yuchen, et al.
Veröffentlicht: (2025)
von: Jiao, Yuchen, et al.
Veröffentlicht: (2025)
A Local Polyak-Lojasiewicz and Descent Lemma of Gradient Descent For Overparametrized Linear Models
von: Xu, Ziqing, et al.
Veröffentlicht: (2025)
von: Xu, Ziqing, et al.
Veröffentlicht: (2025)
DDPM-CD: Denoising Diffusion Probabilistic Models as Feature Extractors for Change Detection
von: Bandara, Wele Gedara Chaminda, et al.
Veröffentlicht: (2022)
von: Bandara, Wele Gedara Chaminda, et al.
Veröffentlicht: (2022)
Transformers with Joint Tokens and Local-Global Attention for Efficient Human Pose Estimation
von: Kinfu, Kaleab A., et al.
Veröffentlicht: (2025)
von: Kinfu, Kaleab A., et al.
Veröffentlicht: (2025)
How Do Nonlinear Transformers Learn and Generalize in In-Context Learning?
von: Li, Hongkang, et al.
Veröffentlicht: (2024)
von: Li, Hongkang, et al.
Veröffentlicht: (2024)
DDPM Score Matching and Distribution Learning
von: Chewi, Sinho, et al.
Veröffentlicht: (2025)
von: Chewi, Sinho, et al.
Veröffentlicht: (2025)
Agnostic Private Density Estimation for GMMs via List Global Stability
von: Afzali, Mohammad, et al.
Veröffentlicht: (2024)
von: Afzali, Mohammad, et al.
Veröffentlicht: (2024)
Optimizing DDPM Sampling with Shortcut Fine-Tuning
von: Fan, Ying, et al.
Veröffentlicht: (2023)
von: Fan, Ying, et al.
Veröffentlicht: (2023)
Understanding the Learning Dynamics of LoRA: A Gradient Flow Perspective on Low-Rank Adaptation in Matrix Factorization
von: Xu, Ziqing, et al.
Veröffentlicht: (2025)
von: Xu, Ziqing, et al.
Veröffentlicht: (2025)
Convergence Rates for Gradient Descent on the Edge of Stability in Overparametrised Least Squares
von: MacDonald, Lachlan Ewen, et al.
Veröffentlicht: (2025)
von: MacDonald, Lachlan Ewen, et al.
Veröffentlicht: (2025)
Learning on Transformers is Provable Low-Rank and Sparse: A One-layer Analysis
von: Li, Hongkang, et al.
Veröffentlicht: (2024)
von: Li, Hongkang, et al.
Veröffentlicht: (2024)
TabDDPM: Modelling Tabular Data with Diffusion Models
von: Kotelnikov, Akim, et al.
Veröffentlicht: (2022)
von: Kotelnikov, Akim, et al.
Veröffentlicht: (2022)
Theoretical Learning Performance of Graph Neural Networks: The Impact of Jumping Connections and Layer-wise Sparsification
von: Sun, Jiawei, et al.
Veröffentlicht: (2025)
von: Sun, Jiawei, et al.
Veröffentlicht: (2025)
Out-of-Distribution Detection in LiDAR Semantic Segmentation Using Epistemic Uncertainty from Hierarchical GMMs
von: Miandashti, Hanieh Shojaei, et al.
Veröffentlicht: (2025)
von: Miandashti, Hanieh Shojaei, et al.
Veröffentlicht: (2025)
Mathematics of Continual Learning
von: Peng, Liangzu, et al.
Veröffentlicht: (2025)
von: Peng, Liangzu, et al.
Veröffentlicht: (2025)
An Edit Friendly DDPM Noise Space: Inversion and Manipulations
von: Huberman-Spiegelglas, Inbar, et al.
Veröffentlicht: (2023)
von: Huberman-Spiegelglas, Inbar, et al.
Veröffentlicht: (2023)
Joint Training of Multi-Token Prediction in Reinforcement Learning via Optimal Coefficient Calibration
von: Wang, Zili, et al.
Veröffentlicht: (2026)
von: Wang, Zili, et al.
Veröffentlicht: (2026)
TransConv-DDPM: Enhanced Diffusion Model for Generating Time-Series Data in Healthcare
von: Kabir, Md Shahriar, et al.
Veröffentlicht: (2026)
von: Kabir, Md Shahriar, et al.
Veröffentlicht: (2026)
How Transformers Learn to Plan via Multi-Token Prediction
von: Huang, Jianhao, et al.
Veröffentlicht: (2026)
von: Huang, Jianhao, et al.
Veröffentlicht: (2026)
Learning safety critics via a non-contractive binary bellman operator
von: Castellano, Agustin, et al.
Veröffentlicht: (2024)
von: Castellano, Agustin, et al.
Veröffentlicht: (2024)
What Improves the Generalization of Graph Transformers? A Theoretical Dive into the Self-attention and Positional Encoding
von: Li, Hongkang, et al.
Veröffentlicht: (2024)
von: Li, Hongkang, et al.
Veröffentlicht: (2024)
When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear Transformers
von: Li, Hongkang, et al.
Veröffentlicht: (2025)
von: Li, Hongkang, et al.
Veröffentlicht: (2025)
Enhancing Graph Transformers with Hierarchical Distance Structural Encoding
von: Luo, Yuankai, et al.
Veröffentlicht: (2023)
von: Luo, Yuankai, et al.
Veröffentlicht: (2023)
Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
von: Li, Hongkang, et al.
Veröffentlicht: (2024)
von: Li, Hongkang, et al.
Veröffentlicht: (2024)
SEMRes-DDPM: Residual Network Based Diffusion Modelling Applied to Imbalanced Data
von: Zheng, Ming, et al.
Veröffentlicht: (2024)
von: Zheng, Ming, et al.
Veröffentlicht: (2024)
Node Identifiers: Compact, Discrete Representations for Efficient Graph Learning
von: Luo, Yuankai, et al.
Veröffentlicht: (2024)
von: Luo, Yuankai, et al.
Veröffentlicht: (2024)
Point-RTD: Replaced Token Denoising for Pretraining Transformer Models on Point Clouds
von: Stone, Gunner, et al.
Veröffentlicht: (2025)
von: Stone, Gunner, et al.
Veröffentlicht: (2025)
Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
von: Li, Hongkang, et al.
Veröffentlicht: (2025)
von: Li, Hongkang, et al.
Veröffentlicht: (2025)
Polyp-DDPM: Diffusion-Based Semantic Polyp Synthesis for Enhanced Segmentation
von: Dorjsembe, Zolnamar, et al.
Veröffentlicht: (2024)
von: Dorjsembe, Zolnamar, et al.
Veröffentlicht: (2024)
Why DDIM Hallucinates More Than DDPM: A Theoretical Analysis of Reverse Dynamics
von: Ashiq, Muhammad H., et al.
Veröffentlicht: (2026)
von: Ashiq, Muhammad H., et al.
Veröffentlicht: (2026)
ECG Signal Denoising Using Multi-scale Patch Embedding and Transformers
von: Zhu, Ding, et al.
Veröffentlicht: (2024)
von: Zhu, Ding, et al.
Veröffentlicht: (2024)
Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning
von: Reuss, Moritz, et al.
Veröffentlicht: (2024)
von: Reuss, Moritz, et al.
Veröffentlicht: (2024)
MedSpaformer: a Transferable Transformer with Multi-granularity Token Sparsification for Medical Time Series Classification
von: Ye, Jiexia, et al.
Veröffentlicht: (2025)
von: Ye, Jiexia, et al.
Veröffentlicht: (2025)
Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion
von: Manor, Hila, et al.
Veröffentlicht: (2024)
von: Manor, Hila, et al.
Veröffentlicht: (2024)
Denoising Diffusions with Optimal Transport: Localization, Curvature, and Multi-Scale Complexity
von: Liang, Tengyuan, et al.
Veröffentlicht: (2024)
von: Liang, Tengyuan, et al.
Veröffentlicht: (2024)
MTM: A Multi-Scale Token Mixing Transformer for Irregular Multivariate Time Series Classification
von: Zhong, Shuhan, et al.
Veröffentlicht: (2025)
von: Zhong, Shuhan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Understanding Incremental Learning with Closed-form Solution to Gradient Flow on Overparamerterized Matrix Factorization
von: Min, Hancheng, et al.
Veröffentlicht: (2025) -
Can Implicit Bias Imply Adversarial Robustness?
von: Min, Hancheng, et al.
Veröffentlicht: (2024) -
Early Neuron Alignment in Two-layer ReLU Networks with Small Initialization
von: Min, Hancheng, et al.
Veröffentlicht: (2023) -
Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data
von: Min, Hancheng, et al.
Veröffentlicht: (2025) -
Optimal Convergence Analysis of DDPM for General Distributions
von: Jiao, Yuchen, et al.
Veröffentlicht: (2025)