Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Mo, Shentong, Tian, Yapeng |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Efficient 3D Shape Generation via Diffusion Mamba with Bidirectional SSMs
par: Mo, Shentong
Publié: (2024)
par: Mo, Shentong
Publié: (2024)
Text-to-Audio Generation Synchronized with Videos
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
Improving Visual Representation Alignment Generation with GRPO
par: Mo, Shentong, et autres
Publié: (2026)
par: Mo, Shentong, et autres
Publié: (2026)
GMAIL: Generative Modality Alignment for generated Image Learning
par: Mo, Shentong, et autres
Publié: (2026)
par: Mo, Shentong, et autres
Publié: (2026)
GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining
par: Mo, Shentong, et autres
Publié: (2026)
par: Mo, Shentong, et autres
Publié: (2026)
Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows
par: Mo, Shentong, et autres
Publié: (2026)
par: Mo, Shentong, et autres
Publié: (2026)
Semantic Grouping Network for Audio Source Separation
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation
par: Mo, Shentong, et autres
Publié: (2026)
par: Mo, Shentong, et autres
Publié: (2026)
pMoE: Prompting Diverse Experts Together Wins More in Visual Adaptation
par: Mo, Shentong, et autres
Publié: (2026)
par: Mo, Shentong, et autres
Publié: (2026)
The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder Learning
par: Mo, Shentong
Publié: (2024)
par: Mo, Shentong
Publié: (2024)
LSPT: Long-term Spatial Prompt Tuning for Visual Representation Learning
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
A Large-scale Medical Visual Task Adaptation Benchmark
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
DMT-JEPA: Discriminative Masked Targets for Joint-Embedding Predictive Architecture
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
MultiMed: Massively Multimodal and Multitask Medical Understanding
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
Modality-Inconsistent Continual Learning of Multimodal Large Language Models
par: Pian, Weiguo, et autres
Publié: (2024)
par: Pian, Weiguo, et autres
Publié: (2024)
TP-Blend: Textual-Prompt Attention Pairing for Precise Object-Style Blending in Diffusion Models
par: Jin, Xin, et autres
Publié: (2026)
par: Jin, Xin, et autres
Publié: (2026)
ViBiDSampler: Enhancing Video Interpolation Using Bidirectional Diffusion Sampler
par: Yang, Serin, et autres
Publié: (2024)
par: Yang, Serin, et autres
Publié: (2024)
Aligning Audio-Visual Joint Representations with an Agentic Workflow
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
Connecting Joint-Embedding Predictive Architecture with Contrastive Self-supervised Learning
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
Attention-Mamba: A Mamba-Enhanced Multi-Scale Parallel Inference Network for Medical Image Segmentation
par: Zhang, Yanhua, et autres
Publié: (2024)
par: Zhang, Yanhua, et autres
Publié: (2024)
DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap
par: Mo, Shentong, et autres
Publié: (2025)
par: Mo, Shentong, et autres
Publié: (2025)
Unified Video-Language Pre-training with Synchronized Audio
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
Scaling Diffusion Transformers Efficiently via $μ$P
par: Zheng, Chenyu, et autres
Publié: (2025)
par: Zheng, Chenyu, et autres
Publié: (2025)
M4V: Multi-Modal Mamba for Text-to-Video Generation
par: Huang, Jiancheng, et autres
Publié: (2025)
par: Huang, Jiancheng, et autres
Publié: (2025)
CLIPErase: Efficient Unlearning of Visual-Textual Associations in CLIP
par: Yang, Tianyu, et autres
Publié: (2024)
par: Yang, Tianyu, et autres
Publié: (2024)
Contextualized Diffusion Models for Text-Guided Image and Video Generation
par: Yang, Ling, et autres
Publié: (2024)
par: Yang, Ling, et autres
Publié: (2024)
Scaling Image and Video Generation via Test-Time Evolutionary Search
par: He, Haoran, et autres
Publié: (2025)
par: He, Haoran, et autres
Publié: (2025)
PoM: Efficient Image and Video Generation with the Polynomial Mixer
par: Picard, David, et autres
Publié: (2024)
par: Picard, David, et autres
Publié: (2024)
Mamba-3D as Masked Autoencoders for Accurate and Data-Efficient Analysis of Medical Ultrasound Videos
par: Zhou, Jiaheng, et autres
Publié: (2025)
par: Zhou, Jiaheng, et autres
Publié: (2025)
End-to-End Multi-Modal Diffusion Mamba
par: Lu, Chunhao, et autres
Publié: (2025)
par: Lu, Chunhao, et autres
Publié: (2025)
IoT-LM: Large Multisensory Language Models for the Internet of Things
par: Mo, Shentong, et autres
Publié: (2024)
par: Mo, Shentong, et autres
Publié: (2024)
MultiIoT: Benchmarking Machine Learning for the Internet of Things
par: Mo, Shentong, et autres
Publié: (2023)
par: Mo, Shentong, et autres
Publié: (2023)
Towards Precise Scaling Laws for Video Diffusion Transformers
par: Yin, Yuanyang, et autres
Publié: (2024)
par: Yin, Yuanyang, et autres
Publié: (2024)
GalaxyDiT: Efficient Video Generation with Guidance Alignment and Adaptive Proxy in Diffusion Transformers
par: Song, Zhiye, et autres
Publié: (2025)
par: Song, Zhiye, et autres
Publié: (2025)
Warped Diffusion: Solving Video Inverse Problems with Image Diffusion Models
par: Daras, Giannis, et autres
Publié: (2024)
par: Daras, Giannis, et autres
Publié: (2024)
VideoDPO: Omni-Preference Alignment for Video Diffusion Generation
par: Liu, Runtao, et autres
Publié: (2024)
par: Liu, Runtao, et autres
Publié: (2024)
From Image to Video: An Empirical Study of Diffusion Representations
par: Vélez, Pedro, et autres
Publié: (2025)
par: Vélez, Pedro, et autres
Publié: (2025)
MambaPEFT: Exploring Parameter-Efficient Fine-Tuning for Mamba
par: Yoshimura, Masakazu, et autres
Publié: (2024)
par: Yoshimura, Masakazu, et autres
Publié: (2024)
Solving Video Inverse Problems Using Image Diffusion Models
par: Kwon, Taesung, et autres
Publié: (2024)
par: Kwon, Taesung, et autres
Publié: (2024)
Conditional Diffusion on Web-Scale Image Pairs leads to Diverse Image Variations
par: Kumar, Manoj, et autres
Publié: (2024)
par: Kumar, Manoj, et autres
Publié: (2024)
Documents similaires
-
Efficient 3D Shape Generation via Diffusion Mamba with Bidirectional SSMs
par: Mo, Shentong
Publié: (2024) -
Text-to-Audio Generation Synchronized with Videos
par: Mo, Shentong, et autres
Publié: (2024) -
Improving Visual Representation Alignment Generation with GRPO
par: Mo, Shentong, et autres
Publié: (2026) -
GMAIL: Generative Modality Alignment for generated Image Learning
par: Mo, Shentong, et autres
Publié: (2026) -
GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining
par: Mo, Shentong, et autres
Publié: (2026)