VideoMamba: State Space Model for Efficient Video Understanding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Kunchang, Li, Xinhao, Wang, Yi, He, Yinan, Wang, Yali, Wang, Limin, Qiao, Yu |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Unmasked Teacher: Towards Training-Efficient Video Foundation Models
par: Li, Kunchang, et autres
Publié: (2023)
par: Li, Kunchang, et autres
Publié: (2023)
Harvest Video Foundation Models via Efficient Post-Pretraining
par: Li, Yizhuo, et autres
Publié: (2023)
par: Li, Yizhuo, et autres
Publié: (2023)
Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
par: Chen, Guo, et autres
Publié: (2024)
par: Chen, Guo, et autres
Publié: (2024)
VideoMamba: Spatio-Temporal Selective State Space Model
par: Park, Jinyoung, et autres
Publié: (2024)
par: Park, Jinyoung, et autres
Publié: (2024)
VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
par: Li, Xinhao, et autres
Publié: (2024)
par: Li, Xinhao, et autres
Publié: (2024)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
par: Li, Kunchang, et autres
Publié: (2023)
par: Li, Kunchang, et autres
Publié: (2023)
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
par: Li, Xinhao, et autres
Publié: (2024)
par: Li, Xinhao, et autres
Publié: (2024)
VideoChat: Chat-Centric Video Understanding
par: Li, KunChang, et autres
Publié: (2023)
par: Li, KunChang, et autres
Publié: (2023)
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
par: Wang, Yi, et autres
Publié: (2024)
par: Wang, Yi, et autres
Publié: (2024)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
par: Wang, Yi, et autres
Publié: (2023)
par: Wang, Yi, et autres
Publié: (2023)
VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
par: Yan, Ziang, et autres
Publié: (2025)
par: Yan, Ziang, et autres
Publié: (2025)
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
par: Yan, Ziang, et autres
Publié: (2024)
par: Yan, Ziang, et autres
Publié: (2024)
Make Your Training Flexible: Towards Deployment-Efficient Video Models
par: Wang, Chenting, et autres
Publié: (2025)
par: Wang, Chenting, et autres
Publié: (2025)
Super Encoding Network: Recursive Association of Multi-Modal Encoders for Video Understanding
par: Chen, Boyu, et autres
Publié: (2025)
par: Chen, Boyu, et autres
Publié: (2025)
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
par: Li, Xinhao, et autres
Publié: (2025)
par: Li, Xinhao, et autres
Publié: (2025)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
par: Zeng, Xiangyu, et autres
Publié: (2024)
par: Zeng, Xiangyu, et autres
Publié: (2024)
Snakes and Ladders: Two Steps Up for VideoMamba
par: Lu, Hui, et autres
Publié: (2024)
par: Lu, Hui, et autres
Publié: (2024)
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
par: Wang, Yi, et autres
Publié: (2025)
par: Wang, Yi, et autres
Publié: (2025)
Percept, Chat, and then Adapt: Multimodal Knowledge Transfer of Foundation Models for Open-World Video Recognition
par: Chen, Boyu, et autres
Publié: (2024)
par: Chen, Boyu, et autres
Publié: (2024)
VideoChat-A1: Thinking with Long Videos by Chain-of-Shot Reasoning
par: Wang, Zikang, et autres
Publié: (2025)
par: Wang, Zikang, et autres
Publié: (2025)
LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering
par: Zhang, Hongjie, et autres
Publié: (2023)
par: Zhang, Hongjie, et autres
Publié: (2023)
Dual Branch VideoMamba with Gated Class Token Fusion for Violence Detection
par: Senadeera, Damith Chamalke, et autres
Publié: (2025)
par: Senadeera, Damith Chamalke, et autres
Publié: (2025)
TransAgent: Transfer Vision-Language Foundation Models with Heterogeneous Agent Collaboration
par: Guo, Yiwei, et autres
Publié: (2024)
par: Guo, Yiwei, et autres
Publié: (2024)
ZeroI2V: Zero-Cost Adaptation of Pre-trained Transformers from Image to Video
par: Li, Xinhao, et autres
Publié: (2023)
par: Li, Xinhao, et autres
Publié: (2023)
StreamForest: Efficient Online Video Understanding with Persistent Event Memory
par: Zeng, Xiangyu, et autres
Publié: (2025)
par: Zeng, Xiangyu, et autres
Publié: (2025)
VideoMAP: Toward Scalable Mamba-based Video Autoregressive Pretraining
par: Liu, Yunze, et autres
Publié: (2025)
par: Liu, Yunze, et autres
Publié: (2025)
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
par: Xu, Yicheng, et autres
Publié: (2025)
par: Xu, Yicheng, et autres
Publié: (2025)
V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents
par: Yue, Zhengrong, et autres
Publié: (2025)
par: Yue, Zhengrong, et autres
Publié: (2025)
Online Video Understanding: OVBench and VideoChat-Online
par: Huang, Zhenpeng, et autres
Publié: (2024)
par: Huang, Zhenpeng, et autres
Publié: (2024)
InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
par: Wang, Chenting, et autres
Publié: (2025)
par: Wang, Chenting, et autres
Publié: (2025)
MambaVF: State Space Model for Efficient Video Fusion
par: Zhao, Zixiang, et autres
Publié: (2026)
par: Zhao, Zixiang, et autres
Publié: (2026)
Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning
par: Zeng, Xiangyu, et autres
Publié: (2026)
par: Zeng, Xiangyu, et autres
Publié: (2026)
EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
par: Pei, Baoqi, et autres
Publié: (2024)
par: Pei, Baoqi, et autres
Publié: (2024)
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
par: Chen, Guo, et autres
Publié: (2024)
par: Chen, Guo, et autres
Publié: (2024)
Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel
par: Wang, Zun, et autres
Publié: (2024)
par: Wang, Zun, et autres
Publié: (2024)
RainMamba: Enhanced Locality Learning with State Space Models for Video Deraining
par: Wu, Hongtao, et autres
Publié: (2024)
par: Wu, Hongtao, et autres
Publié: (2024)
Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning
par: Pei, Baoqi, et autres
Publié: (2025)
par: Pei, Baoqi, et autres
Publié: (2025)
EventMamba: Enhancing Spatio-Temporal Locality with State Space Models for Event-Based Video Reconstruction
par: Ge, Chengjie, et autres
Publié: (2025)
par: Ge, Chengjie, et autres
Publié: (2025)
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
par: Yu, Jiashuo, et autres
Publié: (2025)
par: Yu, Jiashuo, et autres
Publié: (2025)
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
par: Zou, Jialv, et autres
Publié: (2025)
par: Zou, Jialv, et autres
Publié: (2025)
Documents similaires
-
Unmasked Teacher: Towards Training-Efficient Video Foundation Models
par: Li, Kunchang, et autres
Publié: (2023) -
Harvest Video Foundation Models via Efficient Post-Pretraining
par: Li, Yizhuo, et autres
Publié: (2023) -
Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
par: Chen, Guo, et autres
Publié: (2024) -
VideoMamba: Spatio-Temporal Selective State Space Model
par: Park, Jinyoung, et autres
Publié: (2024) -
VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
par: Li, Xinhao, et autres
Publié: (2024)