The Cosine Schedule is Fisher-Rao-Optimal for Masked Discrete Diffusion Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Leo, Syed, Saifuddin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Score-Optimal Diffusion Schedules
von: Williams, Christopher, et al.
Veröffentlicht: (2024)
von: Williams, Christopher, et al.
Veröffentlicht: (2024)
Optimal Inference Schedules for Masked Diffusion Models
von: Chen, Sitan, et al.
Veröffentlicht: (2025)
von: Chen, Sitan, et al.
Veröffentlicht: (2025)
Why Masking Diffusion Works: Condition on the Jump Schedule for Improved Discrete Diffusion
von: Amin, Alan N., et al.
Veröffentlicht: (2025)
von: Amin, Alan N., et al.
Veröffentlicht: (2025)
PPO in the Fisher-Rao geometry
von: Lascu, Razvan-Andrei, et al.
Veröffentlicht: (2025)
von: Lascu, Razvan-Andrei, et al.
Veröffentlicht: (2025)
CREPE: Controlling Diffusion with Replica Exchange
von: He, Jiajun, et al.
Veröffentlicht: (2025)
von: He, Jiajun, et al.
Veröffentlicht: (2025)
Kernel Approximation of Fisher-Rao Gradient Flows
von: Zhu, Jia-Jie, et al.
Veröffentlicht: (2024)
von: Zhu, Jia-Jie, et al.
Veröffentlicht: (2024)
Sampling in Unit Time with Kernel Fisher-Rao Flow
von: Maurais, Aimee, et al.
Veröffentlicht: (2024)
von: Maurais, Aimee, et al.
Veröffentlicht: (2024)
Improved Sampling Schedules for Discrete Diffusion Models
von: Foresti, Alberto, et al.
Veröffentlicht: (2026)
von: Foresti, Alberto, et al.
Veröffentlicht: (2026)
Error Bounds and Optimal Schedules for Masked Diffusions with Factorized Approximations
von: Lavenant, Hugo, et al.
Veröffentlicht: (2025)
von: Lavenant, Hugo, et al.
Veröffentlicht: (2025)
Beyond Masked and Unmasked: Discrete Diffusion Models via Partial Masking
von: Chao, Chen-Hao, et al.
Veröffentlicht: (2025)
von: Chao, Chen-Hao, et al.
Veröffentlicht: (2025)
Functionality-Oriented LLM Merging on the Fisher--Rao Manifold
von: Wang, Jiayu, et al.
Veröffentlicht: (2026)
von: Wang, Jiayu, et al.
Veröffentlicht: (2026)
The Fisher-Rao Loss for Learning under Label Noise
von: Miyamoto, Henrique K., et al.
Veröffentlicht: (2022)
von: Miyamoto, Henrique K., et al.
Veröffentlicht: (2022)
FRInGe: Distribution-Space Integrated Gradients with Fisher--Rao Geometry
von: Martino, Gabriele, et al.
Veröffentlicht: (2026)
von: Martino, Gabriele, et al.
Veröffentlicht: (2026)
An approach to Fisher-Rao metric for infinite dimensional non-parametric information geometry
von: Cheng, Bing, et al.
Veröffentlicht: (2025)
von: Cheng, Bing, et al.
Veröffentlicht: (2025)
Efficient Neural Networks with Discrete Cosine Transform Activations
von: Martinez-Gost, Marc, et al.
Veröffentlicht: (2025)
von: Martinez-Gost, Marc, et al.
Veröffentlicht: (2025)
Simplified and Generalized Masked Diffusion for Discrete Data
von: Shi, Jiaxin, et al.
Veröffentlicht: (2024)
von: Shi, Jiaxin, et al.
Veröffentlicht: (2024)
Multi-Objective Optimization via Wasserstein-Fisher-Rao Gradient Flow
von: Ren, Yinuo, et al.
Veröffentlicht: (2023)
von: Ren, Yinuo, et al.
Veröffentlicht: (2023)
Steering Masked Discrete Diffusion Models via Discrete Denoising Posterior Prediction
von: Rector-Brooks, Jarrid, et al.
Veröffentlicht: (2024)
von: Rector-Brooks, Jarrid, et al.
Veröffentlicht: (2024)
What Exactly Does Guidance Do in Masked Discrete Diffusion Models
von: Ye, He, et al.
Veröffentlicht: (2025)
von: Ye, He, et al.
Veröffentlicht: (2025)
Fisher-Rao distance and pullback SPD cone distances between multivariate normal distributions
von: Nielsen, Frank
Veröffentlicht: (2023)
von: Nielsen, Frank
Veröffentlicht: (2023)
Adaptive function approximation based on the Discrete Cosine Transform (DCT)
von: Pérez-Neira, Ana I., et al.
Veröffentlicht: (2023)
von: Pérez-Neira, Ana I., et al.
Veröffentlicht: (2023)
Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
Latent Shadows: The Gaussian-Discrete Duality in Masked Diffusion
von: Chen, Guinan, et al.
Veröffentlicht: (2026)
von: Chen, Guinan, et al.
Veröffentlicht: (2026)
Weighted Stochastic Differential Equation to Implement Wasserstein-Fisher-Rao Gradient Flow
von: Rahimi, Herlock
Veröffentlicht: (2025)
von: Rahimi, Herlock
Veröffentlicht: (2025)
Inclusive KL Minimization: A Wasserstein-Fisher-Rao Gradient Flow Perspective
von: Zhu, Jia-Jie
Veröffentlicht: (2024)
von: Zhu, Jia-Jie
Veröffentlicht: (2024)
Learning Generation Orders for Masked Discrete Diffusion Models via Variational Inference
von: Fox, David, et al.
Veröffentlicht: (2026)
von: Fox, David, et al.
Veröffentlicht: (2026)
Conditional Diffusion Sampling
von: Castro-Macías, Francisco M., et al.
Veröffentlicht: (2026)
von: Castro-Macías, Francisco M., et al.
Veröffentlicht: (2026)
Accelerated Parallel Tempering via Neural Transports
von: Zhang, Leo, et al.
Veröffentlicht: (2025)
von: Zhang, Leo, et al.
Veröffentlicht: (2025)
Noise Schedule Design for Diffusion Models: An Optimal Control Perspective
von: Kong, Seo Taek, et al.
Veröffentlicht: (2026)
von: Kong, Seo Taek, et al.
Veröffentlicht: (2026)
Boosting Adversarial Training via Fisher-Rao Norm-based Regularization
von: Yin, Xiangyu, et al.
Veröffentlicht: (2024)
von: Yin, Xiangyu, et al.
Veröffentlicht: (2024)
Fisher-Rao Gradient Flow: Geodesic Convexity and Functional Inequalities
von: Carrillo, José A., et al.
Veröffentlicht: (2024)
von: Carrillo, José A., et al.
Veröffentlicht: (2024)
On propagation of chaos for the Fisher-Rao gradient flow in entropic mean-field optimization
von: Lazić, Petra, et al.
Veröffentlicht: (2026)
von: Lazić, Petra, et al.
Veröffentlicht: (2026)
Neural Sampling from Boltzmann Densities: Fisher-Rao Curves in the Wasserstein Geometry
von: Chemseddine, Jannis, et al.
Veröffentlicht: (2024)
von: Chemseddine, Jannis, et al.
Veröffentlicht: (2024)
Fisher Mask Nodes for Language Model Merging
von: K, Thennal D, et al.
Veröffentlicht: (2024)
von: K, Thennal D, et al.
Veröffentlicht: (2024)
Not All Denoising Steps Are Equal: Model Scheduling for Faster Masked Diffusion Language Models
von: Sedykh, Ivan, et al.
Veröffentlicht: (2026)
von: Sedykh, Ivan, et al.
Veröffentlicht: (2026)
Efficient, Multimodal, and Derivative-Free Bayesian Inference With Fisher-Rao Gradient Flows
von: Chen, Yifan, et al.
Veröffentlicht: (2024)
von: Chen, Yifan, et al.
Veröffentlicht: (2024)
FGGM: Fisher-Guided Gradient Masking for Continual Learning
von: Tan, Chao-Hong, et al.
Veröffentlicht: (2026)
von: Tan, Chao-Hong, et al.
Veröffentlicht: (2026)
A Fisher-Rao gradient flow for entropic mean-field min-max games
von: Lascu, Razvan-Andrei, et al.
Veröffentlicht: (2024)
von: Lascu, Razvan-Andrei, et al.
Veröffentlicht: (2024)
Rao-Blackwell Gradient Estimators for Equivariant Denoising Diffusion
von: Tong, Vinh, et al.
Veröffentlicht: (2025)
von: Tong, Vinh, et al.
Veröffentlicht: (2025)
Discrete Cosine Transform Based Decorrelated Attention for Vision Transformers
von: Pan, Hongyi, et al.
Veröffentlicht: (2024)
von: Pan, Hongyi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Score-Optimal Diffusion Schedules
von: Williams, Christopher, et al.
Veröffentlicht: (2024) -
Optimal Inference Schedules for Masked Diffusion Models
von: Chen, Sitan, et al.
Veröffentlicht: (2025) -
Why Masking Diffusion Works: Condition on the Jump Schedule for Improved Discrete Diffusion
von: Amin, Alan N., et al.
Veröffentlicht: (2025) -
PPO in the Fisher-Rao geometry
von: Lascu, Razvan-Andrei, et al.
Veröffentlicht: (2025) -
CREPE: Controlling Diffusion with Replica Exchange
von: He, Jiajun, et al.
Veröffentlicht: (2025)