Any-Order Flexible Length Masked Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Jaeyeon, Cheuk-Kit, Lee, Domingo-Enrich, Carles, Du, Yilun, Kakade, Sham, Ngotiaoco, Timothy, Chen, Sitan, Albergo, Michael |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions
by: Kim, Jaeyeon, et al.
Published: (2025)
by: Kim, Jaeyeon, et al.
Published: (2025)
Selective Underfitting in Diffusion Models
by: Song, Kiwhan, et al.
Published: (2025)
by: Song, Kiwhan, et al.
Published: (2025)
Stop Training for the Worst: Progressive Unmasking Accelerates Masked Diffusion Training
by: Kim, Jaeyeon, et al.
Published: (2026)
by: Kim, Jaeyeon, et al.
Published: (2026)
A unified perspective on fine-tuning and sampling with diffusion and flow models
by: Domingo-Enrich, Carles, et al.
Published: (2026)
by: Domingo-Enrich, Carles, et al.
Published: (2026)
Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models
by: Jelassi, Samy, et al.
Published: (2026)
by: Jelassi, Samy, et al.
Published: (2026)
A Taxonomy of Loss Functions for Stochastic Optimal Control
by: Domingo-Enrich, Carles
Published: (2024)
by: Domingo-Enrich, Carles
Published: (2024)
Tilt Matching for Scalable Sampling and Fine-Tuning
by: Potaptchik, Peter, et al.
Published: (2025)
by: Potaptchik, Peter, et al.
Published: (2025)
Adjoint Matching through the Lens of the Stochastic Maximum Principle in Optimal Control
by: Domingo-Enrich, Carles, et al.
Published: (2026)
by: Domingo-Enrich, Carles, et al.
Published: (2026)
Debiasing Guidance for Discrete Diffusion with Sequential Monte Carlo
by: Lee, Cheuk Kit, et al.
Published: (2025)
by: Lee, Cheuk Kit, et al.
Published: (2025)
No Compute Left Behind: Rethinking Reasoning and Sampling with Masked Diffusion Models
by: Horvitz, Zachary, et al.
Published: (2025)
by: Horvitz, Zachary, et al.
Published: (2025)
Free energy Estimation on Any State Space
by: He, Jiajun, et al.
Published: (2026)
by: He, Jiajun, et al.
Published: (2026)
Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control
by: Domingo-Enrich, Carles, et al.
Published: (2024)
by: Domingo-Enrich, Carles, et al.
Published: (2024)
Compress Then Test: Powerful Kernel Testing in Near-linear Time
by: Domingo-Enrich, Carles, et al.
Published: (2023)
by: Domingo-Enrich, Carles, et al.
Published: (2023)
Test-time scaling of diffusions with flow maps
by: Sabour, Amirmojtaba, et al.
Published: (2025)
by: Sabour, Amirmojtaba, et al.
Published: (2025)
Optimal Inference Schedules for Masked Diffusion Models
by: Chen, Sitan, et al.
Published: (2025)
by: Chen, Sitan, et al.
Published: (2025)
Discrete Tilt Matching
by: Chen, Yuyuan, et al.
Published: (2026)
by: Chen, Yuyuan, et al.
Published: (2026)
Universal Length Generalization with Turing Programs
by: Hou, Kaiying, et al.
Published: (2024)
by: Hou, Kaiying, et al.
Published: (2024)
Conditioning Diffusions Using Malliavin Calculus
by: Pidstrigach, Jakiw, et al.
Published: (2025)
by: Pidstrigach, Jakiw, et al.
Published: (2025)
Cheap Permutation Testing
by: Domingo-Enrich, Carles, et al.
Published: (2025)
by: Domingo-Enrich, Carles, et al.
Published: (2025)
Reinforce Adjoint Matching: Scaling RL Post-Training of Diffusion and Flow-Matching Models
by: Bergmeister, Andreas, et al.
Published: (2026)
by: Bergmeister, Andreas, et al.
Published: (2026)
The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton
by: Abreu, Natalie, et al.
Published: (2025)
by: Abreu, Natalie, et al.
Published: (2025)
The Role of Sparsity for Length Generalization in Transformers
by: Golowich, Noah, et al.
Published: (2025)
by: Golowich, Noah, et al.
Published: (2025)
Neural Optimal Transport with Lagrangian Costs
by: Pooladian, Aram-Alexandre, et al.
Published: (2024)
by: Pooladian, Aram-Alexandre, et al.
Published: (2024)
Rare Event Analysis via Stochastic Optimal Control
by: Du, Yuanqi, et al.
Published: (2026)
by: Du, Yuanqi, et al.
Published: (2026)
Any-Order GPT as Masked Diffusion Model: Decoupling Formulation and Architecture
by: Xue, Shuchen, et al.
Published: (2025)
by: Xue, Shuchen, et al.
Published: (2025)
Trust Region Constrained Measure Transport in Path Space for Stochastic Optimal Control and Inference
by: Blessing, Denis, et al.
Published: (2025)
by: Blessing, Denis, et al.
Published: (2025)
Looped Transformers for Length Generalization
by: Fan, Ying, et al.
Published: (2024)
by: Fan, Ying, et al.
Published: (2024)
Stochastic Optimal Control Matching
by: Domingo-Enrich, Carles, et al.
Published: (2023)
by: Domingo-Enrich, Carles, et al.
Published: (2023)
Value Gradient Guidance for Flow Matching Alignment
by: Liu, Zhen, et al.
Published: (2025)
by: Liu, Zhen, et al.
Published: (2025)
GQ-VAE: A gated quantized VAE for learning variable length tokens
by: Datta, Theo, et al.
Published: (2025)
by: Datta, Theo, et al.
Published: (2025)
Matching the Statistical Query Lower Bound for $k$-Sparse Parity Problems with Sign Stochastic Gradient Descent
by: Kou, Yiwen, et al.
Published: (2024)
by: Kou, Yiwen, et al.
Published: (2024)
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
by: Morwani, Depen, et al.
Published: (2025)
by: Morwani, Depen, et al.
Published: (2025)
Why Depth Matters in Parallelizable Sequence Models: A Lie Algebraic View
by: Heo, Gyuryang, et al.
Published: (2026)
by: Heo, Gyuryang, et al.
Published: (2026)
Prescriptive Scaling Reveals the Evolution of Language Model Capabilities
by: Zhang, Hanlin, et al.
Published: (2026)
by: Zhang, Hanlin, et al.
Published: (2026)
Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
by: Jin, Jikai, et al.
Published: (2025)
by: Jin, Jikai, et al.
Published: (2025)
EvoLM: In Search of Lost Language Model Training Dynamics
by: Qi, Zhenting, et al.
Published: (2025)
by: Qi, Zhenting, et al.
Published: (2025)
Learning Hidden Markov Models Using Conditional Samples
by: Kakade, Sham M., et al.
Published: (2023)
by: Kakade, Sham M., et al.
Published: (2023)
Flash Inference: Near Linear Time Inference for Long Convolution Sequence Models and Beyond
by: Oncescu, Costin-Andrei, et al.
Published: (2024)
by: Oncescu, Costin-Andrei, et al.
Published: (2024)
Learning Flexible Forward Trajectories for Masked Molecular Diffusion
by: Seo, Hyunjin, et al.
Published: (2025)
by: Seo, Hyunjin, et al.
Published: (2025)
Graph Energy Matching: Transport-Aligned Energy-Based Modeling for Graph Generation
by: Balcerak, Michal, et al.
Published: (2026)
by: Balcerak, Michal, et al.
Published: (2026)
Similar Items
-
Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions
by: Kim, Jaeyeon, et al.
Published: (2025) -
Selective Underfitting in Diffusion Models
by: Song, Kiwhan, et al.
Published: (2025) -
Stop Training for the Worst: Progressive Unmasking Accelerates Masked Diffusion Training
by: Kim, Jaeyeon, et al.
Published: (2026) -
A unified perspective on fine-tuning and sampling with diffusion and flow models
by: Domingo-Enrich, Carles, et al.
Published: (2026) -
Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models
by: Jelassi, Samy, et al.
Published: (2026)