The Design Space of Tri-Modal Masked Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Bethune, Louis, Turrisi, Victor, Mlodozeniec, Bruno Kacper, Lopez, Pau Rodriguez, Boominathan, Lokesh, Bhendawade, Nikhil, Shidani, Amitis, Pelemans, Joris, Olausson, Theo X., Hjelm, Devon, Dixon, Paul, Monteiro, Joao, Ablin, Pierre, Banna, Vishnu, Blaas, Arno, Henderson, Nick, Noriy, Kari, Busbridge, Dan, Susskind, Josh, Cuturi, Marco, Belousova, Irina, Zappella, Luca, Webb, Russ, Ramapuram, Jason |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Poly-View Contrastive Learning
by: Shidani, Amitis, et al.
Published: (2024)
by: Shidani, Amitis, et al.
Published: (2024)
Completed Hyperparameter Transfer across Modules, Width, Depth, Batch and Duration
by: Mlodozeniec, Bruno, et al.
Published: (2025)
by: Mlodozeniec, Bruno, et al.
Published: (2025)
Scaling Categorical Flow Maps
by: Davis, Oscar, et al.
Published: (2026)
by: Davis, Oscar, et al.
Published: (2026)
Revisiting the Scaling Properties of Downstream Metrics in Large Language Model Training
by: Krajewski, Jakub, et al.
Published: (2025)
by: Krajewski, Jakub, et al.
Published: (2025)
Learning Unmasking Policies for Diffusion Language Models
by: Jazbec, Metod, et al.
Published: (2025)
by: Jazbec, Metod, et al.
Published: (2025)
Distillation Scaling Laws
by: Busbridge, Dan, et al.
Published: (2025)
by: Busbridge, Dan, et al.
Published: (2025)
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
by: Bethune, Louis, et al.
Published: (2025)
by: Bethune, Louis, et al.
Published: (2025)
Theory, Analysis, and Best Practices for Sigmoid Self-Attention
by: Ramapuram, Jason, et al.
Published: (2024)
by: Ramapuram, Jason, et al.
Published: (2024)
Scaling Properties of Continuous Diffusion Spoken Language Models
by: Ramapuram, Jason, et al.
Published: (2026)
by: Ramapuram, Jason, et al.
Published: (2026)
DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models
by: Monsefi, Amin Karimi, et al.
Published: (2026)
by: Monsefi, Amin Karimi, et al.
Published: (2026)
Controlling Language and Diffusion Models by Transporting Activations
by: Rodriguez, Pau, et al.
Published: (2024)
by: Rodriguez, Pau, et al.
Published: (2024)
LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss
by: Rodriguez, Pau, et al.
Published: (2025)
by: Rodriguez, Pau, et al.
Published: (2025)
Ranking In Generalized Linear Bandits
by: Shidani, Amitis, et al.
Published: (2022)
by: Shidani, Amitis, et al.
Published: (2022)
Interpreting CLIP: Insights on the Robustness to ImageNet Distribution Shifts
by: Crabbé, Jonathan, et al.
Published: (2023)
by: Crabbé, Jonathan, et al.
Published: (2023)
Shielded Diffusion: Generating Novel and Diverse Images using Sparse Repellency
by: Kirchhof, Michael, et al.
Published: (2024)
by: Kirchhof, Michael, et al.
Published: (2024)
DynaMiCS: Fine-tuning LLMs with Performance Constraints using Dynamic Mixtures
by: Gualdoni, Eleonora, et al.
Published: (2026)
by: Gualdoni, Eleonora, et al.
Published: (2026)
M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference
by: Bhendawade, Nikhil, et al.
Published: (2025)
by: Bhendawade, Nikhil, et al.
Published: (2025)
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
by: Saada, Thiziri Nait, et al.
Published: (2025)
by: Saada, Thiziri Nait, et al.
Published: (2025)
Bernstein-type dimension-free concentration for self-normalised martingales
by: Akhavan, Arya, et al.
Published: (2025)
by: Akhavan, Arya, et al.
Published: (2025)
Scaling Laws for Optimal Data Mixtures
by: Shukor, Mustafa, et al.
Published: (2025)
by: Shukor, Mustafa, et al.
Published: (2025)
The Geometries of Truth Are Orthogonal Across Tasks
by: Azizian, Waiss, et al.
Published: (2025)
by: Azizian, Waiss, et al.
Published: (2025)
Sample and Map from a Single Convex Potential: Generation using Conjugate Moment Measures
by: Vesseron, Nina, et al.
Published: (2025)
by: Vesseron, Nina, et al.
Published: (2025)
HyperTransport: Amortized Conditioning of T2I Generative Models
by: Maiorca, Valentino, et al.
Published: (2026)
by: Maiorca, Valentino, et al.
Published: (2026)
Beyond Real Data: Synthetic Data through the Lens of Regularization
by: Shidani, Amitis, et al.
Published: (2025)
by: Shidani, Amitis, et al.
Published: (2025)
GenCtrl -- A Formal Controllability Toolkit for Generative Models
by: Cheng, Emily, et al.
Published: (2026)
by: Cheng, Emily, et al.
Published: (2026)
Amortizing Maximum Inner Product Search with Learned Support Functions
by: Olausson, Theo X., et al.
Published: (2026)
by: Olausson, Theo X., et al.
Published: (2026)
Multivariate Conformal Prediction using Optimal Transport
by: Klein, Michal, et al.
Published: (2025)
by: Klein, Michal, et al.
Published: (2025)
Nectar: Neural Estimation of Cached-Token Attention via Regression
by: Monteiro, João, et al.
Published: (2026)
by: Monteiro, João, et al.
Published: (2026)
Locking Pretrained Weights via Deep Low-Rank Residual Distillation
by: Sakamoto, Keitaro, et al.
Published: (2026)
by: Sakamoto, Keitaro, et al.
Published: (2026)
Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
by: Bhendawade, Nikhil, et al.
Published: (2025)
by: Bhendawade, Nikhil, et al.
Published: (2025)
Speculative Streaming: Fast LLM Inference without Auxiliary Models
by: Bhendawade, Nikhil, et al.
Published: (2024)
by: Bhendawade, Nikhil, et al.
Published: (2024)
FS-DFM: Fast and Accurate Long Text Generation with Few-Step Diffusion Language Models
by: Monsefi, Amin Karimi, et al.
Published: (2025)
by: Monsefi, Amin Karimi, et al.
Published: (2025)
Trajectory as the Teacher: Few-Step Discrete Flow Matching via Energy-Navigated Distillation
by: Monsefi, Amin Karimi, et al.
Published: (2026)
by: Monsefi, Amin Karimi, et al.
Published: (2026)
On the Modeling Capabilities of Large Language Models for Sequential Decision Making
by: Klissarov, Martin, et al.
Published: (2024)
by: Klissarov, Martin, et al.
Published: (2024)
Attention when you need
by: Boominathan, Lokesh, et al.
Published: (2025)
by: Boominathan, Lokesh, et al.
Published: (2025)
Considerations for Distribution Shift Robustness of Diagnostic Models in Healthcare
by: Blaas, Arno, et al.
Published: (2024)
by: Blaas, Arno, et al.
Published: (2024)
GRACE: A Language Model Framework for Explainable Inverse Reinforcement Learning
by: Sapora, Silvia, et al.
Published: (2025)
by: Sapora, Silvia, et al.
Published: (2025)
Peribalus strictus subsp. strictus
by: Belousova, E. N.
Published: (2007)
by: Belousova, E. N.
Published: (2007)
Peribalus (Asioperibalus) przewalskii Belousova 2007, sp. n.
by: Belousova, E. N.
Published: (2007)
by: Belousova, E. N.
Published: (2007)
Status of the fisheries in Rostov Region over the period 2011-2015
by: Belousova, E.V.
Published: (2017)
by: Belousova, E.V.
Published: (2017)
Similar Items
-
Poly-View Contrastive Learning
by: Shidani, Amitis, et al.
Published: (2024) -
Completed Hyperparameter Transfer across Modules, Width, Depth, Batch and Duration
by: Mlodozeniec, Bruno, et al.
Published: (2025) -
Scaling Categorical Flow Maps
by: Davis, Oscar, et al.
Published: (2026) -
Revisiting the Scaling Properties of Downstream Metrics in Large Language Model Training
by: Krajewski, Jakub, et al.
Published: (2025) -
Learning Unmasking Policies for Diffusion Language Models
by: Jazbec, Metod, et al.
Published: (2025)