$\text{PKS}^4$:Parallel Kinematic Selective State Space Scanners for Efficient Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zeng, Lingjie, Zhang, Hailun, Wang, Xiwen, Zhao, Qijun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dehallu3D: Hallucination-Mitigated 3D Generation from Single Image via Cyclic View Consistency Refinement
von: Wang, Xiwen, et al.
Veröffentlicht: (2026)
von: Wang, Xiwen, et al.
Veröffentlicht: (2026)
VideoMamba: State Space Model for Efficient Video Understanding
von: Li, Kunchang, et al.
Veröffentlicht: (2024)
von: Li, Kunchang, et al.
Veröffentlicht: (2024)
FOCUS: Efficient Keyframe Selection for Long Video Understanding
von: Zhu, Zirui, et al.
Veröffentlicht: (2025)
von: Zhu, Zirui, et al.
Veröffentlicht: (2025)
LiquidTAD: Efficient Temporal Action Detection via Parallel Liquid-Inspired Temporal Relaxation
von: Sun, Zepeng, et al.
Veröffentlicht: (2026)
von: Sun, Zepeng, et al.
Veröffentlicht: (2026)
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
von: Zhang, Yunzhu, et al.
Veröffentlicht: (2025)
von: Zhang, Yunzhu, et al.
Veröffentlicht: (2025)
Efficient Self-Supervised Video Hashing with Selective State Spaces
von: Wang, Jinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Jinpeng, et al.
Veröffentlicht: (2024)
Unleashing the Power of Motion and Depth: A Selective Fusion Strategy for RGB-D Video Salient Object Detection
von: He, Jiahao, et al.
Veröffentlicht: (2025)
von: He, Jiahao, et al.
Veröffentlicht: (2025)
MAMBA4D: Efficient Long-Sequence Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space Models
von: Liu, Jiuming, et al.
Veröffentlicht: (2024)
von: Liu, Jiuming, et al.
Veröffentlicht: (2024)
CamoSAM2: Motion-Appearance Induced Auto-Refining Prompts for Video Camouflaged Object Detection
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
von: Chen, Guo, et al.
Veröffentlicht: (2024)
von: Chen, Guo, et al.
Veröffentlicht: (2024)
MambaVF: State Space Model for Efficient Video Fusion
von: Zhao, Zixiang, et al.
Veröffentlicht: (2026)
von: Zhao, Zixiang, et al.
Veröffentlicht: (2026)
On the Perception Bottleneck of VLMs for Chart Understanding
von: Liu, Junteng, et al.
Veröffentlicht: (2025)
von: Liu, Junteng, et al.
Veröffentlicht: (2025)
StreamForest: Efficient Online Video Understanding with Persistent Event Memory
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2025)
Mamba-based Spatio-Frequency Motion Perception for Video Camouflaged Object Detection
von: Li, Xin, et al.
Veröffentlicht: (2025)
von: Li, Xin, et al.
Veröffentlicht: (2025)
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
von: Zou, Jialv, et al.
Veröffentlicht: (2025)
von: Zou, Jialv, et al.
Veröffentlicht: (2025)
Event-Anchored Frame Selection for Effective Long-Video Understanding
von: Chen, Wang, et al.
Veröffentlicht: (2026)
von: Chen, Wang, et al.
Veröffentlicht: (2026)
VideoMamba: Spatio-Temporal Selective State Space Model
von: Park, Jinyoung, et al.
Veröffentlicht: (2024)
von: Park, Jinyoung, et al.
Veröffentlicht: (2024)
Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
von: Chen, Wang, et al.
Veröffentlicht: (2026)
von: Chen, Wang, et al.
Veröffentlicht: (2026)
S3T-Former: A Purely Spike-Driven State-Space Topology Transformer for Skeleton Action Recognition
von: Zheng, Naichuan, et al.
Veröffentlicht: (2026)
von: Zheng, Naichuan, et al.
Veröffentlicht: (2026)
An Efficient Streaming Video Understanding Framework with Agentic Control
von: Liu, Jinming, et al.
Veröffentlicht: (2026)
von: Liu, Jinming, et al.
Veröffentlicht: (2026)
Selective Structured State Space for Multispectral-fused Small Target Detection
von: Zhang, Qianqian, et al.
Veröffentlicht: (2025)
von: Zhang, Qianqian, et al.
Veröffentlicht: (2025)
Adaptively Bypassing Vision Transformer Blocks for Efficient Visual Tracking
von: Yang, Xiangyang, et al.
Veröffentlicht: (2024)
von: Yang, Xiangyang, et al.
Veröffentlicht: (2024)
GIFT: Global Irreplaceability Frame Targeting for Efficient Video Understanding
von: Ma, Junpeng, et al.
Veröffentlicht: (2026)
von: Ma, Junpeng, et al.
Veröffentlicht: (2026)
Trajectory-aware Shifted State Space Models for Online Video Super-Resolution
von: Zhu, Qiang, et al.
Veröffentlicht: (2025)
von: Zhu, Qiang, et al.
Veröffentlicht: (2025)
M-LLM Based Video Frame Selection for Efficient Video Understanding
von: Hu, Kai, et al.
Veröffentlicht: (2025)
von: Hu, Kai, et al.
Veröffentlicht: (2025)
TRAM: Global Trajectory and Motion of 3D Humans from in-the-wild Videos
von: Wang, Yufu, et al.
Veröffentlicht: (2024)
von: Wang, Yufu, et al.
Veröffentlicht: (2024)
KPM-Bench: A Kinematic Parsing Motion Benchmark for Fine-grained Motion-centric Video Understanding
von: Lin, Boda, et al.
Veröffentlicht: (2026)
von: Lin, Boda, et al.
Veröffentlicht: (2026)
On Structured State-Space Duality
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2025)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2025)
State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video Understanding
von: Zhou, Jiahuan, et al.
Veröffentlicht: (2025)
von: Zhou, Jiahuan, et al.
Veröffentlicht: (2025)
From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation
von: Li, Bohan, et al.
Veröffentlicht: (2026)
von: Li, Bohan, et al.
Veröffentlicht: (2026)
Ego4o: Egocentric Human Motion Capture and Understanding from Multi-Modal Input
von: Wang, Jian, et al.
Veröffentlicht: (2025)
von: Wang, Jian, et al.
Veröffentlicht: (2025)
ECMamba: Consolidating Selective State Space Model with Retinex Guidance for Efficient Multiple Exposure Correction
von: Dong, Wei, et al.
Veröffentlicht: (2024)
von: Dong, Wei, et al.
Veröffentlicht: (2024)
Explicit Motion Handling and Interactive Prompting for Video Camouflaged Object Detection
von: Zhang, Xin, et al.
Veröffentlicht: (2024)
von: Zhang, Xin, et al.
Veröffentlicht: (2024)
$\text{S}^{3}$Mamba: Arbitrary-Scale Super-Resolution via Scaleable State Space Model
von: Xia, Peizhe, et al.
Veröffentlicht: (2024)
von: Xia, Peizhe, et al.
Veröffentlicht: (2024)
Towards Long Video Understanding via Fine-detailed Video Story Generation
von: You, Zeng, et al.
Veröffentlicht: (2024)
von: You, Zeng, et al.
Veröffentlicht: (2024)
Towards Training-Free Scene Text Editing
von: Li, Yubo, et al.
Veröffentlicht: (2026)
von: Li, Yubo, et al.
Veröffentlicht: (2026)
Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding
von: Shi, Shuyao, et al.
Veröffentlicht: (2026)
von: Shi, Shuyao, et al.
Veröffentlicht: (2026)
Dynamic Patch-aware Enrichment Transformer for Occluded Person Re-Identification
von: Zhang, Xin, et al.
Veröffentlicht: (2024)
von: Zhang, Xin, et al.
Veröffentlicht: (2024)
Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
StereoDiff: Stereo-Diffusion Synergy for Video Depth Estimation
von: Li, Haodong, et al.
Veröffentlicht: (2025)
von: Li, Haodong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Dehallu3D: Hallucination-Mitigated 3D Generation from Single Image via Cyclic View Consistency Refinement
von: Wang, Xiwen, et al.
Veröffentlicht: (2026) -
VideoMamba: State Space Model for Efficient Video Understanding
von: Li, Kunchang, et al.
Veröffentlicht: (2024) -
FOCUS: Efficient Keyframe Selection for Long Video Understanding
von: Zhu, Zirui, et al.
Veröffentlicht: (2025) -
LiquidTAD: Efficient Temporal Action Detection via Parallel Liquid-Inspired Temporal Relaxation
von: Sun, Zepeng, et al.
Veröffentlicht: (2026) -
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
von: Zhang, Yunzhu, et al.
Veröffentlicht: (2025)