VideoMamba: Spatio-Temporal Selective State Space Model
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Jinyoung, Kim, Hee-Seon, Ko, Kangwook, Kim, Minbeom, Kim, Changick |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Doubly-Universal Adversarial Perturbations: Deceiving Vision-Language Models Across Both Images and Text with a Single Perturbation
by: Kim, Hee-Seon, et al.
Published: (2024)
by: Kim, Hee-Seon, et al.
Published: (2024)
Towards Efficient Vision State Space Models via Token Merging
by: Park, Jinyoung, et al.
Published: (2025)
by: Park, Jinyoung, et al.
Published: (2025)
Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts
by: Kim, Hee-Seon, et al.
Published: (2025)
by: Kim, Hee-Seon, et al.
Published: (2025)
VideoMamba: State Space Model for Efficient Video Understanding
by: Li, Kunchang, et al.
Published: (2024)
by: Li, Kunchang, et al.
Published: (2024)
Snakes and Ladders: Two Steps Up for VideoMamba
by: Lu, Hui, et al.
Published: (2024)
by: Lu, Hui, et al.
Published: (2024)
Spatio-Temporal Proximity-Aware Dual-Path Model for Panoramic Activity Recognition
by: Lee, Sumin, et al.
Published: (2024)
by: Lee, Sumin, et al.
Published: (2024)
Difficulty-aware Balancing Margin Loss for Long-tailed Recognition
by: Son, Minseok, et al.
Published: (2024)
by: Son, Minseok, et al.
Published: (2024)
DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO
by: Park, Jinyoung, et al.
Published: (2025)
by: Park, Jinyoung, et al.
Published: (2025)
EventMamba: Enhancing Spatio-Temporal Locality with State Space Models for Event-Based Video Reconstruction
by: Ge, Chengjie, et al.
Published: (2025)
by: Ge, Chengjie, et al.
Published: (2025)
Prompt Learning via Meta-Regularization
by: Park, Jinyoung, et al.
Published: (2024)
by: Park, Jinyoung, et al.
Published: (2024)
Gather-Scatter Mamba: Accelerating Propagation with Efficient State Space Model
by: Ko, Hyun-kyu, et al.
Published: (2025)
by: Ko, Hyun-kyu, et al.
Published: (2025)
Groupwise Query Specialization and Quality-Aware Multi-Assignment for Transformer-based Visual Relationship Detection
by: Kim, Jongha, et al.
Published: (2024)
by: Kim, Jongha, et al.
Published: (2024)
CNG-SFDA:Clean-and-Noisy Region Guided Online-Offline Source-Free Domain Adaptation
by: Cho, Hyeonwoo, et al.
Published: (2024)
by: Cho, Hyeonwoo, et al.
Published: (2024)
BEEP3D: Box-Supervised End-to-End Pseudo-Mask Generation for 3D Instance Segmentation
by: Yoo, Youngju, et al.
Published: (2025)
by: Yoo, Youngju, et al.
Published: (2025)
Station2Radar: query conditioned gaussian splatting for precipitation field
by: Kim, Doyi, et al.
Published: (2026)
by: Kim, Doyi, et al.
Published: (2026)
Dual Branch VideoMamba with Gated Class Token Fusion for Violence Detection
by: Senadeera, Damith Chamalke, et al.
Published: (2025)
by: Senadeera, Damith Chamalke, et al.
Published: (2025)
SELFI: Selective Fusion of Identity for Generalizable Deepfake Detection
by: Kim, Younghun, et al.
Published: (2025)
by: Kim, Younghun, et al.
Published: (2025)
VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis
by: Park, Jinho, et al.
Published: (2026)
by: Park, Jinho, et al.
Published: (2026)
Bridging Implicit and Explicit Geometric Transformation for Single-Image View Synthesis
by: Park, Byeongjun, et al.
Published: (2022)
by: Park, Byeongjun, et al.
Published: (2022)
Joint-Embedding Predictive Architecture for Self-Supervised Learning of Mask Classification Architecture
by: Kim, Dong-Hee, et al.
Published: (2024)
by: Kim, Dong-Hee, et al.
Published: (2024)
ST-VLM: Kinematic Instruction Tuning for Spatio-Temporal Reasoning in Vision-Language Models
by: Ko, Dohwan, et al.
Published: (2025)
by: Ko, Dohwan, et al.
Published: (2025)
HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning
by: Kim, Minkuk, et al.
Published: (2024)
by: Kim, Minkuk, et al.
Published: (2024)
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
by: Hyun, Jeongseok, et al.
Published: (2025)
by: Hyun, Jeongseok, et al.
Published: (2025)
VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint Modeling
by: Go, Hyojun, et al.
Published: (2025)
by: Go, Hyojun, et al.
Published: (2025)
Flow-Assisted Motion Learning Network for Weakly-Supervised Group Activity Recognition
by: Nugroho, Muhammad Adi, et al.
Published: (2024)
by: Nugroho, Muhammad Adi, et al.
Published: (2024)
Denoising Task Routing for Diffusion Models
by: Park, Byeongjun, et al.
Published: (2023)
by: Park, Byeongjun, et al.
Published: (2023)
SplatFlow: Multi-View Rectified Flow Model for 3D Gaussian Splatting Synthesis
by: Go, Hyojun, et al.
Published: (2024)
by: Go, Hyojun, et al.
Published: (2024)
Similarity-Aware Selective State-Space Modeling for Semantic Correspondence
by: Kim, Seungwook, et al.
Published: (2025)
by: Kim, Seungwook, et al.
Published: (2025)
VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning
by: Lee, Ji Soo, et al.
Published: (2025)
by: Lee, Ji Soo, et al.
Published: (2025)
Latent Space Super-Resolution for Higher-Resolution Image Generation with Diffusion Models
by: Jeong, Jinho, et al.
Published: (2025)
by: Jeong, Jinho, et al.
Published: (2025)
Upsample Anything: A Simple and Hard to Beat Baseline for Feature Upsampling
by: Seo, Minseok, et al.
Published: (2025)
by: Seo, Minseok, et al.
Published: (2025)
Long-tailed Adversarial Training with Self-Distillation
by: Cho, Seungju, et al.
Published: (2025)
by: Cho, Seungju, et al.
Published: (2025)
Indirect Gradient Matching for Adversarial Robust Distillation
by: Lee, Hongsin, et al.
Published: (2023)
by: Lee, Hongsin, et al.
Published: (2023)
Enhancing Robustness in Incremental Learning with Adversarial Training
by: Cho, Seungju, et al.
Published: (2023)
by: Cho, Seungju, et al.
Published: (2023)
Weakly Supervised Video Scene Graph Generation via Natural Language Supervision
by: Kim, Kibum, et al.
Published: (2025)
by: Kim, Kibum, et al.
Published: (2025)
Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval
by: Kim, Minkuk, et al.
Published: (2024)
by: Kim, Minkuk, et al.
Published: (2024)
STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment
by: Lee, Jaewoo, et al.
Published: (2023)
by: Lee, Jaewoo, et al.
Published: (2023)
MoE-GRPO: Optimizing Mixture-of-Experts via Reinforcement Learning in Vision-Language Models
by: Ko, Dohwan, et al.
Published: (2026)
by: Ko, Dohwan, et al.
Published: (2026)
PoseMamba: Monocular 3D Human Pose Estimation with Bidirectional Global-Local Spatio-Temporal State Space Model
by: Huang, Yunlong, et al.
Published: (2024)
by: Huang, Yunlong, et al.
Published: (2024)
Switch Diffusion Transformer: Synergizing Denoising Tasks with Sparse Mixture-of-Experts
by: Park, Byeongjun, et al.
Published: (2024)
by: Park, Byeongjun, et al.
Published: (2024)
Similar Items
-
Doubly-Universal Adversarial Perturbations: Deceiving Vision-Language Models Across Both Images and Text with a Single Perturbation
by: Kim, Hee-Seon, et al.
Published: (2024) -
Towards Efficient Vision State Space Models via Token Merging
by: Park, Jinyoung, et al.
Published: (2025) -
Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts
by: Kim, Hee-Seon, et al.
Published: (2025) -
VideoMamba: State Space Model for Efficient Video Understanding
by: Li, Kunchang, et al.
Published: (2024) -
Snakes and Ladders: Two Steps Up for VideoMamba
by: Lu, Hui, et al.
Published: (2024)