Stop Training for the Worst: Progressive Unmasking Accelerates Masked Diffusion Training
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Jaeyeon, Geuter, Jonathan, Alvarez-Melis, David, Kakade, Sham, Chen, Sitan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions
by: Kim, Jaeyeon, et al.
Published: (2025)
by: Kim, Jaeyeon, et al.
Published: (2025)
Selective Underfitting in Diffusion Models
by: Song, Kiwhan, et al.
Published: (2025)
by: Song, Kiwhan, et al.
Published: (2025)
Any-Order Flexible Length Masked Diffusion
by: Kim, Jaeyeon, et al.
Published: (2025)
by: Kim, Jaeyeon, et al.
Published: (2025)
Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
by: Geuter, Jonathan, et al.
Published: (2025)
by: Geuter, Jonathan, et al.
Published: (2025)
DDEQs: Distributional Deep Equilibrium Models through Wasserstein Gradient Flows
by: Geuter, Jonathan, et al.
Published: (2025)
by: Geuter, Jonathan, et al.
Published: (2025)
Random Scaling of Emergent Capabilities
by: Zhao, Rosie, et al.
Published: (2025)
by: Zhao, Rosie, et al.
Published: (2025)
Optimal Inference Schedules for Masked Diffusion Models
by: Chen, Sitan, et al.
Published: (2025)
by: Chen, Sitan, et al.
Published: (2025)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
by: Meterez, Alexandru, et al.
Published: (2025)
by: Meterez, Alexandru, et al.
Published: (2025)
Characterization and Mitigation of Training Instabilities in Microscaling Formats
by: Su, Huangyuan, et al.
Published: (2025)
by: Su, Huangyuan, et al.
Published: (2025)
Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
by: Liu, Bingbin, et al.
Published: (2025)
by: Liu, Bingbin, et al.
Published: (2025)
Boomerang Distillation Enables Zero-Shot Model Size Interpolation
by: Kangaslahti, Sara, et al.
Published: (2025)
by: Kangaslahti, Sara, et al.
Published: (2025)
Accelerated Sampling from Masked Diffusion Models via Entropy Bounded Unmasking
by: Ben-Hamu, Heli, et al.
Published: (2025)
by: Ben-Hamu, Heli, et al.
Published: (2025)
Understanding and Accelerating the Training of Masked Diffusion Language Models
by: Hong, Chunsan, et al.
Published: (2026)
by: Hong, Chunsan, et al.
Published: (2026)
RoBoN: Routed Online Best-of-n for Test-Time Scaling with Multiple LLMs
by: Geuter, Jonathan, et al.
Published: (2025)
by: Geuter, Jonathan, et al.
Published: (2025)
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
by: Morwani, Depen, et al.
Published: (2025)
by: Morwani, Depen, et al.
Published: (2025)
LOTION: Smoothing the Optimization Landscape for Quantized Training
by: Kwun, Mujin, et al.
Published: (2025)
by: Kwun, Mujin, et al.
Published: (2025)
Beyond Masked and Unmasked: Discrete Diffusion Models via Partial Masking
by: Chao, Chen-Hao, et al.
Published: (2025)
by: Chao, Chen-Hao, et al.
Published: (2025)
GQ-VAE: A gated quantized VAE for learning variable length tokens
by: Datta, Theo, et al.
Published: (2025)
by: Datta, Theo, et al.
Published: (2025)
LoRA Training Provably Converges to a Low-Rank Global Minimum or It Fails Loudly (But it Probably Won't Fail)
by: Kim, Junsu, et al.
Published: (2025)
by: Kim, Junsu, et al.
Published: (2025)
DUEL: Exact Likelihood for Masked Diffusion via Deterministic Unmasking
by: Turok, Gilad, et al.
Published: (2026)
by: Turok, Gilad, et al.
Published: (2026)
Mixture of Parrots: Experts improve memorization more than reasoning
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
Transcendence: Generative Models Can Outperform The Experts That Train Them
by: Zhang, Edwin, et al.
Published: (2024)
by: Zhang, Edwin, et al.
Published: (2024)
Universal Length Generalization with Turing Programs
by: Hou, Kaiying, et al.
Published: (2024)
by: Hou, Kaiying, et al.
Published: (2024)
CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training
by: Brandfonbrener, David, et al.
Published: (2024)
by: Brandfonbrener, David, et al.
Published: (2024)
Matching the Statistical Query Lower Bound for $k$-Sparse Parity Problems with Sign Stochastic Gradient Descent
by: Kou, Yiwen, et al.
Published: (2024)
by: Kou, Yiwen, et al.
Published: (2024)
Soup to go: mitigating forgetting during continual learning with model averaging
by: Kleiman, Anat, et al.
Published: (2025)
by: Kleiman, Anat, et al.
Published: (2025)
EvoLM: In Search of Lost Language Model Training Dynamics
by: Qi, Zhenting, et al.
Published: (2025)
by: Qi, Zhenting, et al.
Published: (2025)
The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton
by: Abreu, Natalie, et al.
Published: (2025)
by: Abreu, Natalie, et al.
Published: (2025)
Consistency Training Helps Stop Sycophancy and Jailbreaks
by: Irpan, Alex, et al.
Published: (2025)
by: Irpan, Alex, et al.
Published: (2025)
Deconstructing What Makes a Good Optimizer for Language Models
by: Zhao, Rosie, et al.
Published: (2024)
by: Zhao, Rosie, et al.
Published: (2024)
Repeat After Me: Transformers are Better than State Space Models at Copying
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
Mask Is What DLLM Needs: A Masked Data Training Paradigm for Diffusion LLMs
by: Ma, Linrui, et al.
Published: (2026)
by: Ma, Linrui, et al.
Published: (2026)
Training-Free Self-Correction for Multimodal Masked Diffusion Models
by: Ouyang, Yidong, et al.
Published: (2026)
by: Ouyang, Yidong, et al.
Published: (2026)
Adversarial Training for Robust Coverage Network under Worst-case Facility Losses
by: Miao, Changhao, et al.
Published: (2026)
by: Miao, Changhao, et al.
Published: (2026)
Bringing Stability to Diffusion: Decomposing and Reducing Variance of Training Masked Diffusion Models
by: Jia, Mengni, et al.
Published: (2025)
by: Jia, Mengni, et al.
Published: (2025)
Q-Probe: A Lightweight Approach to Reward Maximization for Language Models
by: Li, Kenneth, et al.
Published: (2024)
by: Li, Kenneth, et al.
Published: (2024)
Prescriptive Scaling Reveals the Evolution of Language Model Capabilities
by: Zhang, Hanlin, et al.
Published: (2026)
by: Zhang, Hanlin, et al.
Published: (2026)
Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
by: Jin, Jikai, et al.
Published: (2025)
by: Jin, Jikai, et al.
Published: (2025)
Learning Hidden Markov Models Using Conditional Samples
by: Kakade, Sham M., et al.
Published: (2023)
by: Kakade, Sham M., et al.
Published: (2023)
Flash Inference: Near Linear Time Inference for Long Convolution Sequence Models and Beyond
by: Oncescu, Costin-Andrei, et al.
Published: (2024)
by: Oncescu, Costin-Andrei, et al.
Published: (2024)
Similar Items
-
Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions
by: Kim, Jaeyeon, et al.
Published: (2025) -
Selective Underfitting in Diffusion Models
by: Song, Kiwhan, et al.
Published: (2025) -
Any-Order Flexible Length Masked Diffusion
by: Kim, Jaeyeon, et al.
Published: (2025) -
Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
by: Geuter, Jonathan, et al.
Published: (2025) -
DDEQs: Distributional Deep Equilibrium Models through Wasserstein Gradient Flows
by: Geuter, Jonathan, et al.
Published: (2025)