Learning Generation Orders for Masked Discrete Diffusion Models via Variational Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Fox, David, Bowyer, Sam, Liu, Song, Aitchison, Laurence, Santos-Rodriguez, Raul, Yang, Mengyue |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Massively Parallel Expectation Maximization For Approximate Posteriors
by: Heap, Thomas, et al.
Published: (2025)
by: Heap, Thomas, et al.
Published: (2025)
Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints
by: Bowyer, Sam, et al.
Published: (2025)
by: Bowyer, Sam, et al.
Published: (2025)
Why you don't overfit, and don't need Bayes if you only train for one epoch
by: Aitchison, Laurence
Published: (2024)
by: Aitchison, Laurence
Published: (2024)
Learning to Skip the Middle Layers of Transformers
by: Lawson, Tim, et al.
Published: (2025)
by: Lawson, Tim, et al.
Published: (2025)
Controlling changes to attention logits
by: Anson, Ben, et al.
Published: (2025)
by: Anson, Ben, et al.
Published: (2025)
Batch size invariant Adam
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Function-Space Learning Rates
by: Milsom, Edward, et al.
Published: (2025)
by: Milsom, Edward, et al.
Published: (2025)
Using Autodiff to Estimate Posterior Moments, Marginals and Samples
by: Bowyer, Sam, et al.
Published: (2023)
by: Bowyer, Sam, et al.
Published: (2023)
Beyond Masked and Unmasked: Discrete Diffusion Models via Partial Masking
by: Chao, Chen-Hao, et al.
Published: (2025)
by: Chao, Chen-Hao, et al.
Published: (2025)
Inverse-Free Sparse Variational Gaussian Processes
by: Cortinovis, Stefano, et al.
Published: (2026)
by: Cortinovis, Stefano, et al.
Published: (2026)
MONGOOSE: Path-wise Smooth Bayesian Optimisation via Meta-learning
by: Yang, Adam X., et al.
Published: (2023)
by: Yang, Adam X., et al.
Published: (2023)
How to set AdamW's weight decay as you scale model and dataset size
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Simplified and Generalized Masked Diffusion for Discrete Data
by: Shi, Jiaxin, et al.
Published: (2024)
by: Shi, Jiaxin, et al.
Published: (2024)
Steering Masked Discrete Diffusion Models via Discrete Denoising Posterior Prediction
by: Rector-Brooks, Jarrid, et al.
Published: (2024)
by: Rector-Brooks, Jarrid, et al.
Published: (2024)
Masked Diffusion Models are Secretly Learned-Order Autoregressive Models
by: Garg, Prateek, et al.
Published: (2025)
by: Garg, Prateek, et al.
Published: (2025)
Bayesian Low-rank Adaptation for Large Language Models
by: Yang, Adam X., et al.
Published: (2023)
by: Yang, Adam X., et al.
Published: (2023)
Using Neural Networks for Data Cleaning in Weather Datasets
by: Hanslope, Jack R. P., et al.
Published: (2024)
by: Hanslope, Jack R. P., et al.
Published: (2024)
Flexible Infinite-Width Graph Convolutional Neural Networks
by: Anson, Ben, et al.
Published: (2024)
by: Anson, Ben, et al.
Published: (2024)
Convolutional Deep Kernel Machines
by: Milsom, Edward, et al.
Published: (2023)
by: Milsom, Edward, et al.
Published: (2023)
Stochastic Kernel Regularisation Improves Generalisation in Deep Kernel Machines
by: Milsom, Edward, et al.
Published: (2024)
by: Milsom, Edward, et al.
Published: (2024)
Generative Representation Learning on Hyper-relational Knowledge Graphs via Masked Discrete Diffusion
by: Lee, Jaejun, et al.
Published: (2026)
by: Lee, Jaejun, et al.
Published: (2026)
Variational Masked Diffusion Models
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
Optimal Inference Schedules for Masked Diffusion Models
by: Chen, Sitan, et al.
Published: (2025)
by: Chen, Sitan, et al.
Published: (2025)
Scale-invariant Attention
by: Anson, Ben, et al.
Published: (2025)
by: Anson, Ben, et al.
Published: (2025)
Generation Order and Parallel Decoding in Masked Diffusion Models: An Information-Theoretic Perspective
by: Zhang, Shaorong, et al.
Published: (2026)
by: Zhang, Shaorong, et al.
Published: (2026)
DICE: Discrete Inversion Enabling Controllable Editing for Multinomial Diffusion and Masked Generative Models
by: He, Xiaoxiao, et al.
Published: (2024)
by: He, Xiaoxiao, et al.
Published: (2024)
Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers
by: Heap, Thomas, et al.
Published: (2025)
by: Heap, Thomas, et al.
Published: (2025)
Plug-and-Play Controllable Generation for Discrete Masked Models
by: Guo, Wei, et al.
Published: (2024)
by: Guo, Wei, et al.
Published: (2024)
What Exactly Does Guidance Do in Masked Discrete Diffusion Models
by: Ye, He, et al.
Published: (2025)
by: Ye, He, et al.
Published: (2025)
The Cosine Schedule is Fisher-Rao-Optimal for Masked Discrete Diffusion Models
by: Zhang, Leo, et al.
Published: (2025)
by: Zhang, Leo, et al.
Published: (2025)
Masked Completion via Structured Diffusion with White-Box Transformers
by: Pai, Druv, et al.
Published: (2024)
by: Pai, Druv, et al.
Published: (2024)
Variational Transdimensional Inference
by: Davies, Laurence, et al.
Published: (2025)
by: Davies, Laurence, et al.
Published: (2025)
Generative Principal Component Regression via Variational Inference
by: Talbot, Austin, et al.
Published: (2024)
by: Talbot, Austin, et al.
Published: (2024)
Remasking Discrete Diffusion Models with Inference-Time Scaling
by: Wang, Guanghan, et al.
Published: (2025)
by: Wang, Guanghan, et al.
Published: (2025)
Any-Order Flexible Length Masked Diffusion
by: Kim, Jaeyeon, et al.
Published: (2025)
by: Kim, Jaeyeon, et al.
Published: (2025)
KLASS: KL-Guided Fast Inference in Masked Diffusion Models
by: Kim, Seo Hyun, et al.
Published: (2025)
by: Kim, Seo Hyun, et al.
Published: (2025)
Improved Variational Inference in Discrete VAEs using Error Correcting Codes
by: Martínez-García, María, et al.
Published: (2024)
by: Martínez-García, María, et al.
Published: (2024)
Double Descent as a Lens for Sample Efficiency in Autoregressive vs. Discrete Diffusion Models
by: Fraij, Ahmad, et al.
Published: (2025)
by: Fraij, Ahmad, et al.
Published: (2025)
Denoising Diffusion Variational Inference: Diffusion Models as Expressive Variational Posteriors
by: Piriyakulkij, Wasu Top, et al.
Published: (2024)
by: Piriyakulkij, Wasu Top, et al.
Published: (2024)
Efficient Benchmarking Is Just Feature Selection and Multiple Regression
by: Bowyer, Sam, et al.
Published: (2026)
by: Bowyer, Sam, et al.
Published: (2026)
Similar Items
-
Massively Parallel Expectation Maximization For Approximate Posteriors
by: Heap, Thomas, et al.
Published: (2025) -
Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints
by: Bowyer, Sam, et al.
Published: (2025) -
Why you don't overfit, and don't need Bayes if you only train for one epoch
by: Aitchison, Laurence
Published: (2024) -
Learning to Skip the Middle Layers of Transformers
by: Lawson, Tim, et al.
Published: (2025) -
Controlling changes to attention logits
by: Anson, Ben, et al.
Published: (2025)