OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Zou, Jialv, Liao, Bencheng, Zhang, Qian, Liu, Wenyu, Wang, Xinggang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MIM4D: Masked Modeling with Multi-View Video for Autonomous Driving Representation Learning
di: Zou, Jialv, et al.
Pubblicazione: (2024)
di: Zou, Jialv, et al.
Pubblicazione: (2024)
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
di: Zhu, Lianghui, et al.
Pubblicazione: (2024)
di: Zhu, Lianghui, et al.
Pubblicazione: (2024)
Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation
di: Liao, Bencheng, et al.
Pubblicazione: (2025)
di: Liao, Bencheng, et al.
Pubblicazione: (2025)
MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling
di: Li, Yingyue, et al.
Pubblicazione: (2025)
di: Li, Yingyue, et al.
Pubblicazione: (2025)
DiffusionDriveV2: Reinforcement Learning-Constrained Truncated Diffusion Modeling in End-to-End Autonomous Driving
di: Zou, Jialv, et al.
Pubblicazione: (2025)
di: Zou, Jialv, et al.
Pubblicazione: (2025)
OmniMamba4D: Spatio-temporal Mamba for longitudinal CT lesion segmentation
di: Kim, Justin Namuk, et al.
Pubblicazione: (2025)
di: Kim, Justin Namuk, et al.
Pubblicazione: (2025)
InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
di: Tao, Hongyuan, et al.
Pubblicazione: (2025)
di: Tao, Hongyuan, et al.
Pubblicazione: (2025)
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
di: Zeng, Lunbin, et al.
Pubblicazione: (2025)
di: Zeng, Lunbin, et al.
Pubblicazione: (2025)
VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning
di: Jiang, Bo, et al.
Pubblicazione: (2024)
di: Jiang, Bo, et al.
Pubblicazione: (2024)
Lane Graph as Path: Continuity-preserving Path-wise Modeling for Online Lane Graph Construction
di: Liao, Bencheng, et al.
Pubblicazione: (2023)
di: Liao, Bencheng, et al.
Pubblicazione: (2023)
MapTRv2: An End-to-End Framework for Online Vectorized HD Map Construction
di: Liao, Bencheng, et al.
Pubblicazione: (2023)
di: Liao, Bencheng, et al.
Pubblicazione: (2023)
Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
di: Jiang, Bo, et al.
Pubblicazione: (2024)
di: Jiang, Bo, et al.
Pubblicazione: (2024)
ViG: Linear-complexity Visual Sequence Learning with Gated Linear Attention
di: Liao, Bencheng, et al.
Pubblicazione: (2024)
di: Liao, Bencheng, et al.
Pubblicazione: (2024)
ResAD: Normalized Residual Trajectory Modeling for End-to-End Autonomous Driving
di: Zheng, Zhiyu, et al.
Pubblicazione: (2025)
di: Zheng, Zhiyu, et al.
Pubblicazione: (2025)
Omni-Weather: A Unified Multimodal Model for Weather Radar Understanding and Generation
di: Zhou, Zhiwang, et al.
Pubblicazione: (2025)
di: Zhou, Zhiwang, et al.
Pubblicazione: (2025)
UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
di: Li, Yongkang, et al.
Pubblicazione: (2026)
di: Li, Yongkang, et al.
Pubblicazione: (2026)
VideoMamba: State Space Model for Efficient Video Understanding
di: Li, Kunchang, et al.
Pubblicazione: (2024)
di: Li, Kunchang, et al.
Pubblicazione: (2024)
Gait Recognition via Collaborating Discriminative and Generative Diffusion Models
di: Xiong, Haijun, et al.
Pubblicazione: (2025)
di: Xiong, Haijun, et al.
Pubblicazione: (2025)
MambaMIC: An Efficient Baseline for Microscopic Image Classification with State Space Models
di: Zou, Shun, et al.
Pubblicazione: (2024)
di: Zou, Shun, et al.
Pubblicazione: (2024)
Cross-Layer Attentive Feature Upsampling for Low-latency Semantic Segmentation
di: Cheng, Tianheng, et al.
Pubblicazione: (2026)
di: Cheng, Tianheng, et al.
Pubblicazione: (2026)
GaraMoSt: Parallel Multi-Granularity Motion and Structural Modeling for Efficient Multi-Frame Interpolation in DSA Images
di: Xu, Ziyang, et al.
Pubblicazione: (2024)
di: Xu, Ziyang, et al.
Pubblicazione: (2024)
DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention
di: Zhu, Lianghui, et al.
Pubblicazione: (2024)
di: Zhu, Lianghui, et al.
Pubblicazione: (2024)
OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
di: Wang, Zehan, et al.
Pubblicazione: (2024)
di: Wang, Zehan, et al.
Pubblicazione: (2024)
Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
di: Li, Lijiang, et al.
Pubblicazione: (2026)
di: Li, Lijiang, et al.
Pubblicazione: (2026)
AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
di: Jiang, Bo, et al.
Pubblicazione: (2025)
di: Jiang, Bo, et al.
Pubblicazione: (2025)
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
di: Xiao, Yicheng, et al.
Pubblicazione: (2025)
di: Xiao, Yicheng, et al.
Pubblicazione: (2025)
PersonViT: Large-scale Self-supervised Vision Transformer for Person Re-Identification
di: Hu, Bin, et al.
Pubblicazione: (2024)
di: Hu, Bin, et al.
Pubblicazione: (2024)
X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction
di: Ren, Xiaoming, et al.
Pubblicazione: (2026)
di: Ren, Xiaoming, et al.
Pubblicazione: (2026)
DeltaMIL: Gated Memory Integration for Efficient and Discriminative Whole Slide Image Analysis
di: Zhu, Yueting, et al.
Pubblicazione: (2025)
di: Zhu, Yueting, et al.
Pubblicazione: (2025)
VL-Mamba: Exploring State Space Models for Multimodal Learning
di: Qiao, Yanyuan, et al.
Pubblicazione: (2024)
di: Qiao, Yanyuan, et al.
Pubblicazione: (2024)
OmniGen: Unified Multimodal Sensor Generation for Autonomous Driving
di: Tang, Tao, et al.
Pubblicazione: (2025)
di: Tang, Tao, et al.
Pubblicazione: (2025)
Matte Anything: Interactive Natural Image Matting with Segment Anything Models
di: Yao, Jingfeng, et al.
Pubblicazione: (2023)
di: Yao, Jingfeng, et al.
Pubblicazione: (2023)
RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning
di: Gao, Hao, et al.
Pubblicazione: (2025)
di: Gao, Hao, et al.
Pubblicazione: (2025)
UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding
di: Xu, Chenkai, et al.
Pubblicazione: (2025)
di: Xu, Chenkai, et al.
Pubblicazione: (2025)
EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
di: He, Xin, et al.
Pubblicazione: (2025)
di: He, Xin, et al.
Pubblicazione: (2025)
Omni$^2$: Unifying Omnidirectional Image Generation and Editing in an Omni Model
di: Yang, Liu, et al.
Pubblicazione: (2025)
di: Yang, Liu, et al.
Pubblicazione: (2025)
MolSight: Optical Chemical Structure Recognition with SMILES Pretraining, Multi-Granularity Learning and Reinforcement Learning
di: Zhang, Wenrui, et al.
Pubblicazione: (2025)
di: Zhang, Wenrui, et al.
Pubblicazione: (2025)
Occupancy as Set of Points
di: Shi, Yiang, et al.
Pubblicazione: (2024)
di: Shi, Yiang, et al.
Pubblicazione: (2024)
Senna-2: Aligning VLM and End-to-End Driving Policy for Consistent Decision Making and Planning
di: Song, Yuehao, et al.
Pubblicazione: (2026)
di: Song, Yuehao, et al.
Pubblicazione: (2026)
Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing
di: Liu, Jialun, et al.
Pubblicazione: (2026)
di: Liu, Jialun, et al.
Pubblicazione: (2026)
Documenti analoghi
-
MIM4D: Masked Modeling with Multi-View Video for Autonomous Driving Representation Learning
di: Zou, Jialv, et al.
Pubblicazione: (2024) -
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
di: Zhu, Lianghui, et al.
Pubblicazione: (2024) -
Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation
di: Liao, Bencheng, et al.
Pubblicazione: (2025) -
MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling
di: Li, Yingyue, et al.
Pubblicazione: (2025) -
DiffusionDriveV2: Reinforcement Learning-Constrained Truncated Diffusion Modeling in End-to-End Autonomous Driving
di: Zou, Jialv, et al.
Pubblicazione: (2025)