The Cosine Schedule is Fisher-Rao-Optimal for Masked Discrete Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Leo, Syed, Saifuddin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Score-Optimal Diffusion Schedules
by: Williams, Christopher, et al.
Published: (2024)
by: Williams, Christopher, et al.
Published: (2024)
Optimal Inference Schedules for Masked Diffusion Models
by: Chen, Sitan, et al.
Published: (2025)
by: Chen, Sitan, et al.
Published: (2025)
Why Masking Diffusion Works: Condition on the Jump Schedule for Improved Discrete Diffusion
by: Amin, Alan N., et al.
Published: (2025)
by: Amin, Alan N., et al.
Published: (2025)
PPO in the Fisher-Rao geometry
by: Lascu, Razvan-Andrei, et al.
Published: (2025)
by: Lascu, Razvan-Andrei, et al.
Published: (2025)
CREPE: Controlling Diffusion with Replica Exchange
by: He, Jiajun, et al.
Published: (2025)
by: He, Jiajun, et al.
Published: (2025)
Kernel Approximation of Fisher-Rao Gradient Flows
by: Zhu, Jia-Jie, et al.
Published: (2024)
by: Zhu, Jia-Jie, et al.
Published: (2024)
Sampling in Unit Time with Kernel Fisher-Rao Flow
by: Maurais, Aimee, et al.
Published: (2024)
by: Maurais, Aimee, et al.
Published: (2024)
Improved Sampling Schedules for Discrete Diffusion Models
by: Foresti, Alberto, et al.
Published: (2026)
by: Foresti, Alberto, et al.
Published: (2026)
Error Bounds and Optimal Schedules for Masked Diffusions with Factorized Approximations
by: Lavenant, Hugo, et al.
Published: (2025)
by: Lavenant, Hugo, et al.
Published: (2025)
Beyond Masked and Unmasked: Discrete Diffusion Models via Partial Masking
by: Chao, Chen-Hao, et al.
Published: (2025)
by: Chao, Chen-Hao, et al.
Published: (2025)
Functionality-Oriented LLM Merging on the Fisher--Rao Manifold
by: Wang, Jiayu, et al.
Published: (2026)
by: Wang, Jiayu, et al.
Published: (2026)
The Fisher-Rao Loss for Learning under Label Noise
by: Miyamoto, Henrique K., et al.
Published: (2022)
by: Miyamoto, Henrique K., et al.
Published: (2022)
FRInGe: Distribution-Space Integrated Gradients with Fisher--Rao Geometry
by: Martino, Gabriele, et al.
Published: (2026)
by: Martino, Gabriele, et al.
Published: (2026)
An approach to Fisher-Rao metric for infinite dimensional non-parametric information geometry
by: Cheng, Bing, et al.
Published: (2025)
by: Cheng, Bing, et al.
Published: (2025)
Efficient Neural Networks with Discrete Cosine Transform Activations
by: Martinez-Gost, Marc, et al.
Published: (2025)
by: Martinez-Gost, Marc, et al.
Published: (2025)
Simplified and Generalized Masked Diffusion for Discrete Data
by: Shi, Jiaxin, et al.
Published: (2024)
by: Shi, Jiaxin, et al.
Published: (2024)
Multi-Objective Optimization via Wasserstein-Fisher-Rao Gradient Flow
by: Ren, Yinuo, et al.
Published: (2023)
by: Ren, Yinuo, et al.
Published: (2023)
Steering Masked Discrete Diffusion Models via Discrete Denoising Posterior Prediction
by: Rector-Brooks, Jarrid, et al.
Published: (2024)
by: Rector-Brooks, Jarrid, et al.
Published: (2024)
What Exactly Does Guidance Do in Masked Discrete Diffusion Models
by: Ye, He, et al.
Published: (2025)
by: Ye, He, et al.
Published: (2025)
Fisher-Rao distance and pullback SPD cone distances between multivariate normal distributions
by: Nielsen, Frank
Published: (2023)
by: Nielsen, Frank
Published: (2023)
Adaptive function approximation based on the Discrete Cosine Transform (DCT)
by: Pérez-Neira, Ana I., et al.
Published: (2023)
by: Pérez-Neira, Ana I., et al.
Published: (2023)
Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Latent Shadows: The Gaussian-Discrete Duality in Masked Diffusion
by: Chen, Guinan, et al.
Published: (2026)
by: Chen, Guinan, et al.
Published: (2026)
Weighted Stochastic Differential Equation to Implement Wasserstein-Fisher-Rao Gradient Flow
by: Rahimi, Herlock
Published: (2025)
by: Rahimi, Herlock
Published: (2025)
Inclusive KL Minimization: A Wasserstein-Fisher-Rao Gradient Flow Perspective
by: Zhu, Jia-Jie
Published: (2024)
by: Zhu, Jia-Jie
Published: (2024)
Learning Generation Orders for Masked Discrete Diffusion Models via Variational Inference
by: Fox, David, et al.
Published: (2026)
by: Fox, David, et al.
Published: (2026)
Conditional Diffusion Sampling
by: Castro-Macías, Francisco M., et al.
Published: (2026)
by: Castro-Macías, Francisco M., et al.
Published: (2026)
Accelerated Parallel Tempering via Neural Transports
by: Zhang, Leo, et al.
Published: (2025)
by: Zhang, Leo, et al.
Published: (2025)
Noise Schedule Design for Diffusion Models: An Optimal Control Perspective
by: Kong, Seo Taek, et al.
Published: (2026)
by: Kong, Seo Taek, et al.
Published: (2026)
Boosting Adversarial Training via Fisher-Rao Norm-based Regularization
by: Yin, Xiangyu, et al.
Published: (2024)
by: Yin, Xiangyu, et al.
Published: (2024)
Fisher-Rao Gradient Flow: Geodesic Convexity and Functional Inequalities
by: Carrillo, José A., et al.
Published: (2024)
by: Carrillo, José A., et al.
Published: (2024)
On propagation of chaos for the Fisher-Rao gradient flow in entropic mean-field optimization
by: Lazić, Petra, et al.
Published: (2026)
by: Lazić, Petra, et al.
Published: (2026)
Neural Sampling from Boltzmann Densities: Fisher-Rao Curves in the Wasserstein Geometry
by: Chemseddine, Jannis, et al.
Published: (2024)
by: Chemseddine, Jannis, et al.
Published: (2024)
Fisher Mask Nodes for Language Model Merging
by: K, Thennal D, et al.
Published: (2024)
by: K, Thennal D, et al.
Published: (2024)
Not All Denoising Steps Are Equal: Model Scheduling for Faster Masked Diffusion Language Models
by: Sedykh, Ivan, et al.
Published: (2026)
by: Sedykh, Ivan, et al.
Published: (2026)
Efficient, Multimodal, and Derivative-Free Bayesian Inference With Fisher-Rao Gradient Flows
by: Chen, Yifan, et al.
Published: (2024)
by: Chen, Yifan, et al.
Published: (2024)
FGGM: Fisher-Guided Gradient Masking for Continual Learning
by: Tan, Chao-Hong, et al.
Published: (2026)
by: Tan, Chao-Hong, et al.
Published: (2026)
A Fisher-Rao gradient flow for entropic mean-field min-max games
by: Lascu, Razvan-Andrei, et al.
Published: (2024)
by: Lascu, Razvan-Andrei, et al.
Published: (2024)
Rao-Blackwell Gradient Estimators for Equivariant Denoising Diffusion
by: Tong, Vinh, et al.
Published: (2025)
by: Tong, Vinh, et al.
Published: (2025)
Discrete Cosine Transform Based Decorrelated Attention for Vision Transformers
by: Pan, Hongyi, et al.
Published: (2024)
by: Pan, Hongyi, et al.
Published: (2024)
Similar Items
-
Score-Optimal Diffusion Schedules
by: Williams, Christopher, et al.
Published: (2024) -
Optimal Inference Schedules for Masked Diffusion Models
by: Chen, Sitan, et al.
Published: (2025) -
Why Masking Diffusion Works: Condition on the Jump Schedule for Improved Discrete Diffusion
by: Amin, Alan N., et al.
Published: (2025) -
PPO in the Fisher-Rao geometry
by: Lascu, Razvan-Andrei, et al.
Published: (2025) -
CREPE: Controlling Diffusion with Replica Exchange
by: He, Jiajun, et al.
Published: (2025)