Efficient 3D Shape Generation via Diffusion Mamba with Bidirectional SSMs
Fuente:
arXiv
Salvato in:
| Autore principale: | Mo, Shentong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation
di: Mo, Shentong, et al.
Pubblicazione: (2024)
di: Mo, Shentong, et al.
Pubblicazione: (2024)
Improving Visual Representation Alignment Generation with GRPO
di: Mo, Shentong, et al.
Pubblicazione: (2026)
di: Mo, Shentong, et al.
Pubblicazione: (2026)
GMAIL: Generative Modality Alignment for generated Image Learning
di: Mo, Shentong, et al.
Pubblicazione: (2026)
di: Mo, Shentong, et al.
Pubblicazione: (2026)
LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation
di: Mo, Shentong, et al.
Pubblicazione: (2026)
di: Mo, Shentong, et al.
Pubblicazione: (2026)
pMoE: Prompting Diverse Experts Together Wins More in Visual Adaptation
di: Mo, Shentong, et al.
Pubblicazione: (2026)
di: Mo, Shentong, et al.
Pubblicazione: (2026)
The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder Learning
di: Mo, Shentong
Pubblicazione: (2024)
di: Mo, Shentong
Pubblicazione: (2024)
Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows
di: Mo, Shentong, et al.
Pubblicazione: (2026)
di: Mo, Shentong, et al.
Pubblicazione: (2026)
LSPT: Long-term Spatial Prompt Tuning for Visual Representation Learning
di: Mo, Shentong, et al.
Pubblicazione: (2024)
di: Mo, Shentong, et al.
Pubblicazione: (2024)
A Large-scale Medical Visual Task Adaptation Benchmark
di: Mo, Shentong, et al.
Pubblicazione: (2024)
di: Mo, Shentong, et al.
Pubblicazione: (2024)
DMT-JEPA: Discriminative Masked Targets for Joint-Embedding Predictive Architecture
di: Mo, Shentong, et al.
Pubblicazione: (2024)
di: Mo, Shentong, et al.
Pubblicazione: (2024)
MultiMed: Massively Multimodal and Multitask Medical Understanding
di: Mo, Shentong, et al.
Pubblicazione: (2024)
di: Mo, Shentong, et al.
Pubblicazione: (2024)
GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining
di: Mo, Shentong, et al.
Pubblicazione: (2026)
di: Mo, Shentong, et al.
Pubblicazione: (2026)
Text-to-Audio Generation Synchronized with Videos
di: Mo, Shentong, et al.
Pubblicazione: (2024)
di: Mo, Shentong, et al.
Pubblicazione: (2024)
Aligning Audio-Visual Joint Representations with an Agentic Workflow
di: Mo, Shentong, et al.
Pubblicazione: (2024)
di: Mo, Shentong, et al.
Pubblicazione: (2024)
Connecting Joint-Embedding Predictive Architecture with Contrastive Self-supervised Learning
di: Mo, Shentong, et al.
Pubblicazione: (2024)
di: Mo, Shentong, et al.
Pubblicazione: (2024)
DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap
di: Mo, Shentong, et al.
Pubblicazione: (2025)
di: Mo, Shentong, et al.
Pubblicazione: (2025)
Large Intestine 3D Shape Refinement Using Point Diffusion Models for Digital Phantom Generation
di: Mouheb, Kaouther, et al.
Pubblicazione: (2023)
di: Mouheb, Kaouther, et al.
Pubblicazione: (2023)
Mamba-3D as Masked Autoencoders for Accurate and Data-Efficient Analysis of Medical Ultrasound Videos
di: Zhou, Jiaheng, et al.
Pubblicazione: (2025)
di: Zhou, Jiaheng, et al.
Pubblicazione: (2025)
Mamba3D: Enhancing Local Features for 3D Point Cloud Analysis via State Space Model
di: Han, Xu, et al.
Pubblicazione: (2024)
di: Han, Xu, et al.
Pubblicazione: (2024)
End-to-End Multi-Modal Diffusion Mamba
di: Lu, Chunhao, et al.
Pubblicazione: (2025)
di: Lu, Chunhao, et al.
Pubblicazione: (2025)
IoT-LM: Large Multisensory Language Models for the Internet of Things
di: Mo, Shentong, et al.
Pubblicazione: (2024)
di: Mo, Shentong, et al.
Pubblicazione: (2024)
MultiIoT: Benchmarking Machine Learning for the Internet of Things
di: Mo, Shentong, et al.
Pubblicazione: (2023)
di: Mo, Shentong, et al.
Pubblicazione: (2023)
MambaPEFT: Exploring Parameter-Efficient Fine-Tuning for Mamba
di: Yoshimura, Masakazu, et al.
Pubblicazione: (2024)
di: Yoshimura, Masakazu, et al.
Pubblicazione: (2024)
ViBiDSampler: Enhancing Video Interpolation Using Bidirectional Diffusion Sampler
di: Yang, Serin, et al.
Pubblicazione: (2024)
di: Yang, Serin, et al.
Pubblicazione: (2024)
IM-3D: Iterative Multiview Diffusion and Reconstruction for High-Quality 3D Generation
di: Melas-Kyriazi, Luke, et al.
Pubblicazione: (2024)
di: Melas-Kyriazi, Luke, et al.
Pubblicazione: (2024)
RadMamba: Efficient Human Activity Recognition through Radar-based Micro-Doppler-Oriented Mamba State-Space Model
di: Wu, Yizhuo, et al.
Pubblicazione: (2025)
di: Wu, Yizhuo, et al.
Pubblicazione: (2025)
Rethinking Metrics and Diffusion Architecture for 3D Point Cloud Generation
di: Bastico, Matteo, et al.
Pubblicazione: (2025)
di: Bastico, Matteo, et al.
Pubblicazione: (2025)
DF-Mamba: Deformable State Space Modeling for 3D Hand Pose Estimation in Interactions
di: Zhou, Yifan, et al.
Pubblicazione: (2025)
di: Zhou, Yifan, et al.
Pubblicazione: (2025)
VCMamba: Bridging Convolutions with Multi-Directional Mamba for Efficient Visual Representation
di: Munir, Mustafa, et al.
Pubblicazione: (2025)
di: Munir, Mustafa, et al.
Pubblicazione: (2025)
ProxT2I: Efficient Reward-Guided Text-to-Image Generation via Proximal Diffusion
di: Fang, Zhenghan, et al.
Pubblicazione: (2025)
di: Fang, Zhenghan, et al.
Pubblicazione: (2025)
MambaOut: Do We Really Need Mamba for Vision?
di: Yu, Weihao, et al.
Pubblicazione: (2024)
di: Yu, Weihao, et al.
Pubblicazione: (2024)
Shape-Guided Diffusion with Inside-Outside Attention
di: Park, Dong Huk, et al.
Pubblicazione: (2022)
di: Park, Dong Huk, et al.
Pubblicazione: (2022)
Bridging the Gap between 2D and 3D Visual Question Answering: A Fusion Approach for 3D VQA
di: Mo, Wentao, et al.
Pubblicazione: (2024)
di: Mo, Wentao, et al.
Pubblicazione: (2024)
Unified Video-Language Pre-training with Synchronized Audio
di: Mo, Shentong, et al.
Pubblicazione: (2024)
di: Mo, Shentong, et al.
Pubblicazione: (2024)
Geometric Point Attention Transformer for 3D Shape Reassembly
di: Li, Jiahan, et al.
Pubblicazione: (2024)
di: Li, Jiahan, et al.
Pubblicazione: (2024)
Scaling Diffusion Transformers Efficiently via $μ$P
di: Zheng, Chenyu, et al.
Pubblicazione: (2025)
di: Zheng, Chenyu, et al.
Pubblicazione: (2025)
MambaMixer: Efficient Selective State Space Models with Dual Token and Channel Selection
di: Behrouz, Ali, et al.
Pubblicazione: (2024)
di: Behrouz, Ali, et al.
Pubblicazione: (2024)
M4V: Multi-Modal Mamba for Text-to-Video Generation
di: Huang, Jiancheng, et al.
Pubblicazione: (2025)
di: Huang, Jiancheng, et al.
Pubblicazione: (2025)
HiFA: High-fidelity Text-to-3D Generation with Advanced Diffusion Guidance
di: Zhu, Junzhe, et al.
Pubblicazione: (2023)
di: Zhu, Junzhe, et al.
Pubblicazione: (2023)
SkelMamba: A State Space Model for Efficient Skeleton Action Recognition of Neurological Disorders
di: Martinel, Niki, et al.
Pubblicazione: (2024)
di: Martinel, Niki, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation
di: Mo, Shentong, et al.
Pubblicazione: (2024) -
Improving Visual Representation Alignment Generation with GRPO
di: Mo, Shentong, et al.
Pubblicazione: (2026) -
GMAIL: Generative Modality Alignment for generated Image Learning
di: Mo, Shentong, et al.
Pubblicazione: (2026) -
LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation
di: Mo, Shentong, et al.
Pubblicazione: (2026) -
pMoE: Prompting Diverse Experts Together Wins More in Visual Adaptation
di: Mo, Shentong, et al.
Pubblicazione: (2026)