Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Liao, Bencheng, Tao, Hongyuan, Zhang, Qian, Cheng, Tianheng, Li, Yingyue, Yin, Haoran, Liu, Wenyu, Wang, Xinggang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
by: Zou, Jialv, et al.
Published: (2025)
by: Zou, Jialv, et al.
Published: (2025)
MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling
by: Li, Yingyue, et al.
Published: (2025)
by: Li, Yingyue, et al.
Published: (2025)
InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
by: Tao, Hongyuan, et al.
Published: (2025)
by: Tao, Hongyuan, et al.
Published: (2025)
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
by: Zhu, Lianghui, et al.
Published: (2024)
by: Zhu, Lianghui, et al.
Published: (2024)
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
by: Zeng, Lunbin, et al.
Published: (2025)
by: Zeng, Lunbin, et al.
Published: (2025)
Lane Graph as Path: Continuity-preserving Path-wise Modeling for Online Lane Graph Construction
by: Liao, Bencheng, et al.
Published: (2023)
by: Liao, Bencheng, et al.
Published: (2023)
MIM4D: Masked Modeling with Multi-View Video for Autonomous Driving Representation Learning
by: Zou, Jialv, et al.
Published: (2024)
by: Zou, Jialv, et al.
Published: (2024)
Cross-Layer Attentive Feature Upsampling for Low-latency Semantic Segmentation
by: Cheng, Tianheng, et al.
Published: (2026)
by: Cheng, Tianheng, et al.
Published: (2026)
ViG: Linear-complexity Visual Sequence Learning with Gated Linear Attention
by: Liao, Bencheng, et al.
Published: (2024)
by: Liao, Bencheng, et al.
Published: (2024)
Occupancy as Set of Points
by: Shi, Yiang, et al.
Published: (2024)
by: Shi, Yiang, et al.
Published: (2024)
VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning
by: Jiang, Bo, et al.
Published: (2024)
by: Jiang, Bo, et al.
Published: (2024)
Polar Parametrization for Vision-based Surround-View 3D Detection
by: Chen, Shaoyu, et al.
Published: (2022)
by: Chen, Shaoyu, et al.
Published: (2022)
Mask-Adapter: The Devil is in the Masks for Open-Vocabulary Segmentation
by: Li, Yongkang, et al.
Published: (2024)
by: Li, Yongkang, et al.
Published: (2024)
Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
by: Jiang, Bo, et al.
Published: (2024)
by: Jiang, Bo, et al.
Published: (2024)
DiffusionDriveV2: Reinforcement Learning-Constrained Truncated Diffusion Modeling in End-to-End Autonomous Driving
by: Zou, Jialv, et al.
Published: (2025)
by: Zou, Jialv, et al.
Published: (2025)
MapTRv2: An End-to-End Framework for Online Vectorized HD Map Construction
by: Liao, Bencheng, et al.
Published: (2023)
by: Liao, Bencheng, et al.
Published: (2023)
YOLO-World: Real-Time Open-Vocabulary Object Detection
by: Cheng, Tianheng, et al.
Published: (2024)
by: Cheng, Tianheng, et al.
Published: (2024)
VL-Mamba: Exploring State Space Models for Multimodal Learning
by: Qiao, Yanyuan, et al.
Published: (2024)
by: Qiao, Yanyuan, et al.
Published: (2024)
DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention
by: Zhu, Lianghui, et al.
Published: (2024)
by: Zhu, Lianghui, et al.
Published: (2024)
DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving
by: Liao, Bencheng, et al.
Published: (2024)
by: Liao, Bencheng, et al.
Published: (2024)
RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning
by: Gao, Hao, et al.
Published: (2025)
by: Gao, Hao, et al.
Published: (2025)
GaussTR: Foundation Model-Aligned Gaussian Transformer for Self-Supervised 3D Spatial Understanding
by: Jiang, Haoyi, et al.
Published: (2024)
by: Jiang, Haoyi, et al.
Published: (2024)
WeakTr: Exploring Plain Vision Transformer for Weakly-supervised Semantic Segmentation
by: Zhu, Lianghui, et al.
Published: (2023)
by: Zhu, Lianghui, et al.
Published: (2023)
MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking
by: Liu, Xinqi, et al.
Published: (2024)
by: Liu, Xinqi, et al.
Published: (2024)
ControlAR: Controllable Image Generation with Autoregressive Models
by: Li, Zongming, et al.
Published: (2024)
by: Li, Zongming, et al.
Published: (2024)
CrossLLM-Mamba: Multimodal State Space Fusion of LLMs for RNA Interaction Prediction
by: Sadia, Rabeya Tus, et al.
Published: (2026)
by: Sadia, Rabeya Tus, et al.
Published: (2026)
MambaFusion: Adaptive State-Space Fusion for Multimodal 3D Object Detection
by: Narayanan, Venkatraman, et al.
Published: (2026)
by: Narayanan, Venkatraman, et al.
Published: (2026)
Multimodal Dataset Distillation via Phased Teacher Models
by: Guo, Shengbin, et al.
Published: (2026)
by: Guo, Shengbin, et al.
Published: (2026)
FusionMamba: Dynamic Feature Enhancement for Multimodal Image Fusion with Mamba
by: Xie, Xinyu, et al.
Published: (2024)
by: Xie, Xinyu, et al.
Published: (2024)
EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model
by: Zhang, Yuxuan, et al.
Published: (2024)
by: Zhang, Yuxuan, et al.
Published: (2024)
AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
by: Jiang, Bo, et al.
Published: (2025)
by: Jiang, Bo, et al.
Published: (2025)
CardiacMamba: A Multimodal RGB-RF Fusion Framework with State Space Models for Remote Physiological Measurement
by: Wu, Zheng, et al.
Published: (2025)
by: Wu, Zheng, et al.
Published: (2025)
FasterDiT: Towards Faster Diffusion Transformers Training without Architecture Modification
by: Yao, Jingfeng, et al.
Published: (2024)
by: Yao, Jingfeng, et al.
Published: (2024)
LocalMamba: Visual State Space Model with Windowed Selective Scan
by: Huang, Tao, et al.
Published: (2024)
by: Huang, Tao, et al.
Published: (2024)
Gait Recognition via Collaborating Discriminative and Generative Diffusion Models
by: Xiong, Haijun, et al.
Published: (2025)
by: Xiong, Haijun, et al.
Published: (2025)
DefMamba: Deformable Visual State Space Model
by: Liu, Leiye, et al.
Published: (2025)
by: Liu, Leiye, et al.
Published: (2025)
Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
by: Li, Jiaze, et al.
Published: (2026)
by: Li, Jiaze, et al.
Published: (2026)
Point Cloud Mamba: Point Cloud Learning via State Space Model
by: Zhang, Tao, et al.
Published: (2024)
by: Zhang, Tao, et al.
Published: (2024)
PersonViT: Large-scale Self-supervised Vision Transformer for Person Re-Identification
by: Hu, Bin, et al.
Published: (2024)
by: Hu, Bin, et al.
Published: (2024)
Decoding Visual Neural Representations by Multimodal with Dynamic Balancing
by: sun, Kaili, et al.
Published: (2025)
by: sun, Kaili, et al.
Published: (2025)
Similar Items
-
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
by: Zou, Jialv, et al.
Published: (2025) -
MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling
by: Li, Yingyue, et al.
Published: (2025) -
InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
by: Tao, Hongyuan, et al.
Published: (2025) -
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
by: Zhu, Lianghui, et al.
Published: (2024) -
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
by: Zeng, Lunbin, et al.
Published: (2025)