DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone
Fuente:
arXiv
Saved in:
| Main Authors: | Singh, Vaibhav, Ostapenko, Oleksiy, Noël, Pierre-André, Belilovsky, Eugene, Scholak, Torsten |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Model Parallelism With Subnetwork Data Parallelism
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training
by: Ostapenko, Oleksiy, et al.
Published: (2025)
by: Ostapenko, Oleksiy, et al.
Published: (2025)
Unifying Autoregressive and Diffusion-Based Sequence Generation
by: Fathi, Nima, et al.
Published: (2025)
by: Fathi, Nima, et al.
Published: (2025)
Decision Mamba Architectures
by: Correia, André, et al.
Published: (2024)
by: Correia, André, et al.
Published: (2024)
Mamba Modulation: On the Length Generalization of Mamba
by: Lu, Peng, et al.
Published: (2025)
by: Lu, Peng, et al.
Published: (2025)
Apriel-H1: Towards Efficient Enterprise Reasoning Models
by: Ostapenko, Oleksiy, et al.
Published: (2025)
by: Ostapenko, Oleksiy, et al.
Published: (2025)
PI-Mamba: Linear-Time Protein Backbone Generation via Spectrally Initialized Flow Matching
by: Wu, Tianyu, et al.
Published: (2026)
by: Wu, Tianyu, et al.
Published: (2026)
DeciMamba: Exploring the Length Extrapolation Potential of Mamba
by: Ben-Kish, Assaf, et al.
Published: (2024)
by: Ben-Kish, Assaf, et al.
Published: (2024)
Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
by: Legate, Gwen, et al.
Published: (2025)
by: Legate, Gwen, et al.
Published: (2025)
Celo2: Towards Learned Optimization Free Lunch
by: Moudgil, Abhinav, et al.
Published: (2026)
by: Moudgil, Abhinav, et al.
Published: (2026)
Efficient Refusal Ablation in LLM through Optimal Transport
by: Nanfack, Geraldin, et al.
Published: (2026)
by: Nanfack, Geraldin, et al.
Published: (2026)
ms-Mamba: Multi-scale Mamba for Time-Series Forecasting
by: Karadag, Yusuf Meric, et al.
Published: (2025)
by: Karadag, Yusuf Meric, et al.
Published: (2025)
InfoMamba: An Attention-Free Hybrid Mamba-Transformer Model
by: Wang, Youjin, et al.
Published: (2026)
by: Wang, Youjin, et al.
Published: (2026)
eMamba: Efficient Acceleration Framework for Mamba Models in Edge Computing
by: Kim, Jiyong, et al.
Published: (2025)
by: Kim, Jiyong, et al.
Published: (2025)
MambaSL: Exploring Single-Layer Mamba for Time Series Classification
by: Jung, Yoo-Min, et al.
Published: (2026)
by: Jung, Yoo-Min, et al.
Published: (2026)
Don't Retrain, Align: Adapting Autoregressive LMs to Diffusion LMs via Representation Alignment
by: Peng, Fred Zhangzhi, et al.
Published: (2026)
by: Peng, Fred Zhangzhi, et al.
Published: (2026)
Not Only the Last-Layer Features for Spurious Correlations: All Layer Deep Feature Reweighting
by: Hameed, Humza Wajid, et al.
Published: (2024)
by: Hameed, Humza Wajid, et al.
Published: (2024)
A Survey of Mamba
by: Qu, Haohao, et al.
Published: (2024)
by: Qu, Haohao, et al.
Published: (2024)
BioMamba: Leveraging Spectro-Temporal Embedding in Bidirectional Mamba for Enhanced Biosignal Classification
by: Qian, Jian, et al.
Published: (2025)
by: Qian, Jian, et al.
Published: (2025)
ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training
by: Nabli, Adel, et al.
Published: (2024)
by: Nabli, Adel, et al.
Published: (2024)
Differential Mamba
by: Schneider, Nadav, et al.
Published: (2025)
by: Schneider, Nadav, et al.
Published: (2025)
MedMamba: Recasting Mamba for Medical Time Series Classification
by: He, ZhengXiao, et al.
Published: (2026)
by: He, ZhengXiao, et al.
Published: (2026)
When Data Falls Short: Grokking Below the Critical Threshold
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Dual-Phase Continual Learning: Supervised Adaptation Meets Unsupervised Retention
by: Singh, Vaibhav, et al.
Published: (2024)
by: Singh, Vaibhav, et al.
Published: (2024)
Less is More: Undertraining Experts Improves Model Upcycling
by: Horoi, Stefan, et al.
Published: (2025)
by: Horoi, Stefan, et al.
Published: (2025)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
by: Xu, Zukang, et al.
Published: (2025)
by: Xu, Zukang, et al.
Published: (2025)
MamTiff-CAD: Multi-Scale Latent Diffusion with Mamba+ for Complex Parametric Sequence
by: Deng, Liyuan, et al.
Published: (2025)
by: Deng, Liyuan, et al.
Published: (2025)
MambaNet: Mamba-assisted Channel Estimation Neural Network With Attention Mechanism
by: Luan, Dianxin, et al.
Published: (2026)
by: Luan, Dianxin, et al.
Published: (2026)
Guiding Language Model Reasoning with Planning Tokens
by: Wang, Xinyi, et al.
Published: (2023)
by: Wang, Xinyi, et al.
Published: (2023)
End-to-End Multi-Modal Diffusion Mamba
by: Lu, Chunhao, et al.
Published: (2025)
by: Lu, Chunhao, et al.
Published: (2025)
Non-Uniform Parameter-Wise Model Merging
by: Camacho, Albert Manuel Orozco, et al.
Published: (2024)
by: Camacho, Albert Manuel Orozco, et al.
Published: (2024)
The Mamba in the Llama: Distilling and Accelerating Hybrid Models
by: Wang, Junxiong, et al.
Published: (2024)
by: Wang, Junxiong, et al.
Published: (2024)
LE-PDE++: Mamba for accelerating PDEs Simulations
by: Liang, Aoming, et al.
Published: (2024)
by: Liang, Aoming, et al.
Published: (2024)
MambaOut: Do We Really Need Mamba for Vision?
by: Yu, Weihao, et al.
Published: (2024)
by: Yu, Weihao, et al.
Published: (2024)
MambaCapsule: Towards Transparent Cardiac Disease Diagnosis with Electrocardiography Using Mamba Capsule Network
by: Xu, Yinlong, et al.
Published: (2024)
by: Xu, Yinlong, et al.
Published: (2024)
Block-Biased Mamba for Long-Range Sequence Processing
by: Yu, Annan, et al.
Published: (2025)
by: Yu, Annan, et al.
Published: (2025)
DMamba: Decomposition-enhanced Mamba for Time Series Forecasting
by: Chen, Ruxuan, et al.
Published: (2026)
by: Chen, Ruxuan, et al.
Published: (2026)
Mamba State-Space Models Are Lyapunov-Stable Learners
by: Halloran, John T., et al.
Published: (2024)
by: Halloran, John T., et al.
Published: (2024)
Sequential Order-Robust Mamba for Time Series Forecasting
by: Lee, Seunghan, et al.
Published: (2024)
by: Lee, Seunghan, et al.
Published: (2024)
A Mamba Foundation Model for Time Series Forecasting
by: Ma, Haoyu, et al.
Published: (2024)
by: Ma, Haoyu, et al.
Published: (2024)
Similar Items
-
Model Parallelism With Subnetwork Data Parallelism
by: Singh, Vaibhav, et al.
Published: (2025) -
Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training
by: Ostapenko, Oleksiy, et al.
Published: (2025) -
Unifying Autoregressive and Diffusion-Based Sequence Generation
by: Fathi, Nima, et al.
Published: (2025) -
Decision Mamba Architectures
by: Correia, André, et al.
Published: (2024) -
Mamba Modulation: On the Length Generalization of Mamba
by: Lu, Peng, et al.
Published: (2025)