DiMSUM: Diffusion Mamba -- A Scalable and Unified Spatial-Frequency Method for Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Phung, Hao, Dao, Quan, Dao, Trung, Phan, Hoang, Metaxas, Dimitris, Tran, Anh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Text-to-Image Generation
by: Dao, Quan, et al.
Published: (2024)
by: Dao, Quan, et al.
Published: (2024)
MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model
by: Dao, Quan, et al.
Published: (2026)
by: Dao, Quan, et al.
Published: (2026)
An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental Learning
by: Tran, Quyen, et al.
Published: (2022)
by: Tran, Quyen, et al.
Published: (2022)
Improved Training Technique for Latent Consistency Models
by: Dao, Quan, et al.
Published: (2025)
by: Dao, Quan, et al.
Published: (2025)
Supercharged One-step Text-to-Image Diffusion Models with Negative Prompts
by: Nguyen, Viet, et al.
Published: (2024)
by: Nguyen, Viet, et al.
Published: (2024)
A High-Quality Robust Diffusion Framework for Corrupted Dataset
by: Dao, Quan, et al.
Published: (2023)
by: Dao, Quan, et al.
Published: (2023)
AutoEdit: Automatic Hyperparameter Tuning for Image Editing
by: Pham, Chau, et al.
Published: (2025)
by: Pham, Chau, et al.
Published: (2025)
Generalization Bounds for Robust Contrastive Learning: From Theory to Practice
by: Tran, Ngoc N., et al.
Published: (2023)
by: Tran, Ngoc N., et al.
Published: (2023)
Improved Training Technique for Shortcut Models
by: Nguyen, Anh, et al.
Published: (2025)
by: Nguyen, Anh, et al.
Published: (2025)
SwiftBrush v2: Make Your One-step Diffusion Model Better Than Its Teacher
by: Dao, Trung, et al.
Published: (2024)
by: Dao, Trung, et al.
Published: (2024)
EFHQ: Multi-purpose ExtremePose-Face-HQ dataset
by: Dao, Trung Tuan, et al.
Published: (2023)
by: Dao, Trung Tuan, et al.
Published: (2023)
Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
by: Dao, Quan, et al.
Published: (2025)
by: Dao, Quan, et al.
Published: (2025)
PrefPaint: Enhancing Medical Image Inpainting through Expert Human Feedback
by: Bui, Duy-Bao, et al.
Published: (2025)
by: Bui, Duy-Bao, et al.
Published: (2025)
Channel-Partitioned Windowed Attention And Frequency Learning for Single Image Super-Resolution
by: Tran, Dinh Phu, et al.
Published: (2024)
by: Tran, Dinh Phu, et al.
Published: (2024)
KOPPA: Improving Prompt-based Continual Learning with Key-Query Orthogonal Projection and Prototype-based One-Versus-All
by: Tran, Quyen, et al.
Published: (2023)
by: Tran, Quyen, et al.
Published: (2023)
VSRM: A Robust Mamba-Based Framework for Video Super-Resolution
by: Tran, Dinh Phu, et al.
Published: (2025)
by: Tran, Dinh Phu, et al.
Published: (2025)
Multi-Level CLS Token Fusion for Contrastive Learning in Endoscopy Image Classification
by: Nguyen, Y Hop, et al.
Published: (2025)
by: Nguyen, Y Hop, et al.
Published: (2025)
SwiftBrush: One-Step Text-to-Image Diffusion Model with Variational Score Distillation
by: Nguyen, Thuan Hoang, et al.
Published: (2023)
by: Nguyen, Thuan Hoang, et al.
Published: (2023)
Enhancing Domain Adaptation through Prompt Gradient Alignment
by: Phan, Hoang, et al.
Published: (2024)
by: Phan, Hoang, et al.
Published: (2024)
Towards Efficient and Robust Moment Retrieval System: A Unified Framework for Multi-Granularity Models and Temporal Reranking
by: Tran, Huu-Loc, et al.
Published: (2025)
by: Tran, Huu-Loc, et al.
Published: (2025)
MasHeNe: A Benchmark for Head and Neck CT Mass Segmentation using Window-Enhanced Mamba with Frequency-Domain Integration
by: Dao, Thao Thi Phuong, et al.
Published: (2025)
by: Dao, Thao Thi Phuong, et al.
Published: (2025)
Score-Guided Diffusion for 3D Human Recovery
by: Stathopoulos, Anastasis, et al.
Published: (2024)
by: Stathopoulos, Anastasis, et al.
Published: (2024)
Preserving Clusters in Prompt Learning for Unsupervised Domain Adaptation
by: Vuong, Tung-Long, et al.
Published: (2025)
by: Vuong, Tung-Long, et al.
Published: (2025)
Frequency Attention for Knowledge Distillation
by: Pham, Cuong, et al.
Published: (2024)
by: Pham, Cuong, et al.
Published: (2024)
Hiding and Recovering Knowledge in Text-to-Image Diffusion Models via Learnable Prompts
by: Bui, Anh, et al.
Published: (2024)
by: Bui, Anh, et al.
Published: (2024)
Conditional Diffusion Model for Longitudinal Medical Image Generation
by: Dao, Duy-Phuong, et al.
Published: (2024)
by: Dao, Duy-Phuong, et al.
Published: (2024)
SINE: SINgle Image Editing with Text-to-Image Diffusion Models
by: Zhang, Zhixing, et al.
Published: (2022)
by: Zhang, Zhixing, et al.
Published: (2022)
NAYER: Noisy Layer Data Generation for Efficient and Effective Data-free Knowledge Distillation
by: Tran, Minh-Tuan, et al.
Published: (2023)
by: Tran, Minh-Tuan, et al.
Published: (2023)
Beyond Motion Pattern: An Empirical Study of Physical Forces for Human Motion Understanding
by: Dao, Anh, et al.
Published: (2025)
by: Dao, Anh, et al.
Published: (2025)
DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis
by: Teng, Yao, et al.
Published: (2024)
by: Teng, Yao, et al.
Published: (2024)
Leveraging Lightweight Entity Extraction for Scalable Event-Based Image Retrieval
by: Minh, Dao Sy Duy, et al.
Published: (2025)
by: Minh, Dao Sy Duy, et al.
Published: (2025)
GriDiT: Factorized Grid-Based Diffusion for Efficient Long Image Sequence Generation
by: Tomar, Snehal Singh, et al.
Published: (2025)
by: Tomar, Snehal Singh, et al.
Published: (2025)
Scalable Autoregressive Image Generation with Mamba
by: Li, Haopeng, et al.
Published: (2024)
by: Li, Haopeng, et al.
Published: (2024)
CLIMB: Controllable Longitudinal Brain Image Generation using Mamba-based Latent Diffusion Model and Gaussian-aligned Autoencoder
by: Dao, Duy-Phuong, et al.
Published: (2026)
by: Dao, Duy-Phuong, et al.
Published: (2026)
Cycle Training with Semi-Supervised Domain Adaptation: Bridging Accuracy and Efficiency for Real-Time Mobile Scene Detection
by: Phan-Nguyen, Huu-Phong, et al.
Published: (2025)
by: Phan-Nguyen, Huu-Phong, et al.
Published: (2025)
Spatial-Frequency Enhanced Mamba for Multi-Modal Image Fusion
by: Sun, Hui, et al.
Published: (2025)
by: Sun, Hui, et al.
Published: (2025)
Test-Time Instance-Specific Parameter Composition: A New Paradigm for Adaptive Generative Modeling
by: Tran, Minh-Tuan, et al.
Published: (2026)
by: Tran, Minh-Tuan, et al.
Published: (2026)
Overcoming the Curvature Bottleneck in MeanFlow
by: Zhang, Xinxi, et al.
Published: (2025)
by: Zhang, Xinxi, et al.
Published: (2025)
Steering Rectified Flow Models in the Vector Field for Controlled Image Generation
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
LoG-VMamba: Local-Global Vision Mamba for Medical Image Segmentation
by: Dang, Trung Dinh Quoc, et al.
Published: (2024)
by: Dang, Trung Dinh Quoc, et al.
Published: (2024)
Similar Items
-
Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Text-to-Image Generation
by: Dao, Quan, et al.
Published: (2024) -
MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model
by: Dao, Quan, et al.
Published: (2026) -
An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental Learning
by: Tran, Quyen, et al.
Published: (2022) -
Improved Training Technique for Latent Consistency Models
by: Dao, Quan, et al.
Published: (2025) -
Supercharged One-step Text-to-Image Diffusion Models with Negative Prompts
by: Nguyen, Viet, et al.
Published: (2024)