HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
Fuente:
arXiv
Saved in:
| Main Authors: | An, Joungbin, Grauman, Kristen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding
by: An, Joungbin, et al.
Published: (2026)
by: An, Joungbin, et al.
Published: (2026)
Progress-Aware Video Frame Captioning
by: Xue, Zihui, et al.
Published: (2024)
by: Xue, Zihui, et al.
Published: (2024)
Learning Skill-Attributes for Transferable Assessment in Video
by: Ashutosh, Kumar, et al.
Published: (2025)
by: Ashutosh, Kumar, et al.
Published: (2025)
ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos
by: Somayazulu, Arjun, et al.
Published: (2026)
by: Somayazulu, Arjun, et al.
Published: (2026)
Mamba-VMR: Multimodal Query Augmentation via Generated Videos for Precise Temporal Grounding
by: Sun, Yunzhuo, et al.
Published: (2026)
by: Sun, Yunzhuo, et al.
Published: (2026)
SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
Mamba-VGGT: Persistent Long-Sequence Video Geometry Grounded Transformer via External Sliding Window Mamba Memory
by: Deng, Tianchen, et al.
Published: (2026)
by: Deng, Tianchen, et al.
Published: (2026)
When Thinking Drifts: Evidential Grounding for Robust Video Reasoning
by: Luo, Mi, et al.
Published: (2025)
by: Luo, Mi, et al.
Published: (2025)
Hi-Mamba: Hierarchical Mamba for Efficient Image Super-Resolution
by: Qiao, Junbo, et al.
Published: (2024)
by: Qiao, Junbo, et al.
Published: (2024)
HieraVid: Hierarchical Token Pruning for Fast Video Large Language Models
by: Guo, Yansong, et al.
Published: (2026)
by: Guo, Yansong, et al.
Published: (2026)
FIction: 4D Future Interaction Prediction from Video
by: Ashutosh, Kumar, et al.
Published: (2024)
by: Ashutosh, Kumar, et al.
Published: (2024)
Learning Object State Changes in Videos: An Open-World Perspective
by: Xue, Zihui, et al.
Published: (2023)
by: Xue, Zihui, et al.
Published: (2023)
Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video
by: Adebi, Daniel, et al.
Published: (2025)
by: Adebi, Daniel, et al.
Published: (2025)
Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models
by: Baid, Ami, et al.
Published: (2026)
by: Baid, Ami, et al.
Published: (2026)
VideoMamba: Spatio-Temporal Selective State Space Model
by: Park, Jinyoung, et al.
Published: (2024)
by: Park, Jinyoung, et al.
Published: (2024)
EndoMamba: An Efficient Foundation Model for Endoscopic Videos via Hierarchical Pre-training
by: Tian, Qingyao, et al.
Published: (2025)
by: Tian, Qingyao, et al.
Published: (2025)
Stitch-a-Demo: Video Demonstrations from Multistep Descriptions
by: Wu, Chi Hsuan, et al.
Published: (2025)
by: Wu, Chi Hsuan, et al.
Published: (2025)
AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation
by: Gong, Sitong, et al.
Published: (2025)
by: Gong, Sitong, et al.
Published: (2025)
EgoExo-WM: Unlocking Exo Video for Ego World Models
by: Tran, Danny, et al.
Published: (2026)
by: Tran, Danny, et al.
Published: (2026)
SportSkills: Physical Skill Learning from Sports Instructional Videos
by: Ashutosh, Kumar, et al.
Published: (2026)
by: Ashutosh, Kumar, et al.
Published: (2026)
Detours for Navigating Instructional Videos
by: Ashutosh, Kumar, et al.
Published: (2024)
by: Ashutosh, Kumar, et al.
Published: (2024)
Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
PlainMamba: Improving Non-Hierarchical Mamba in Visual Recognition
by: Yang, Chenhongyi, et al.
Published: (2024)
by: Yang, Chenhongyi, et al.
Published: (2024)
InterMamba: Efficient Human-Human Interaction Generation with Adaptive Spatio-Temporal Mamba
by: Wu, Zizhao, et al.
Published: (2025)
by: Wu, Zizhao, et al.
Published: (2025)
PhysMamba: Efficient Remote Physiological Measurement with SlowFast Temporal Difference Mamba
by: Luo, Chaoqi, et al.
Published: (2024)
by: Luo, Chaoqi, et al.
Published: (2024)
Mamba-in-Mamba: Centralized Mamba-Cross-Scan in Tokenized Mamba Model for Hyperspectral Image Classification
by: Zhou, Weilian, et al.
Published: (2024)
by: Zhou, Weilian, et al.
Published: (2024)
Text-controlled Motion Mamba: Text-Instructed Temporal Grounding of Human Motion
by: Wang, Xinghan, et al.
Published: (2024)
by: Wang, Xinghan, et al.
Published: (2024)
HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness
by: Xue, Zihui, et al.
Published: (2024)
by: Xue, Zihui, et al.
Published: (2024)
FluencyVE: Marrying Temporal-Aware Mamba with Bypass Attention for Video Editing
by: Cai, Mingshu, et al.
Published: (2025)
by: Cai, Mingshu, et al.
Published: (2025)
MamFusion: Multi-Mamba with Temporal Fusion for Partially Relevant Video Retrieval
by: Ying, Xinru, et al.
Published: (2025)
by: Ying, Xinru, et al.
Published: (2025)
MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed Videos
by: Sinha, Arkaprava, et al.
Published: (2025)
by: Sinha, Arkaprava, et al.
Published: (2025)
STNMamba: Mamba-based Spatial-Temporal Normality Learning for Video Anomaly Detection
by: Li, Zhangxun, et al.
Published: (2024)
by: Li, Zhangxun, et al.
Published: (2024)
MLVTG: Mamba-Based Feature Alignment and LLM-Driven Purification for Multi-Modal Video Temporal Grounding
by: Zhu, Zhiyi, et al.
Published: (2025)
by: Zhu, Zhiyi, et al.
Published: (2025)
Put Myself in Your Shoes: Lifting the Egocentric Perspective from Exocentric Videos
by: Luo, Mi, et al.
Published: (2024)
by: Luo, Mi, et al.
Published: (2024)
Matten: Video Generation with Mamba-Attention
by: Gao, Yu, et al.
Published: (2024)
by: Gao, Yu, et al.
Published: (2024)
Fast Vision Mamba: Pooling Spatial Dimensions for Accelerated Processing
by: Kapse, Saarthak, et al.
Published: (2025)
by: Kapse, Saarthak, et al.
Published: (2025)
Learning Human Motion with Temporally Conditional Mamba
by: Nguyen, Quang, et al.
Published: (2025)
by: Nguyen, Quang, et al.
Published: (2025)
HieraSurg: Hierarchy-Aware Diffusion Model for Surgical Video Generation
by: Biagini, Diego, et al.
Published: (2025)
by: Biagini, Diego, et al.
Published: (2025)
MistExit: Learning to Exit for Early Mistake Detection in Procedural Videos
by: Majumder, Sagnik, et al.
Published: (2026)
by: Majumder, Sagnik, et al.
Published: (2026)
Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos
by: Majumder, Sagnik, et al.
Published: (2024)
by: Majumder, Sagnik, et al.
Published: (2024)
Similar Items
-
UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding
by: An, Joungbin, et al.
Published: (2026) -
Progress-Aware Video Frame Captioning
by: Xue, Zihui, et al.
Published: (2024) -
Learning Skill-Attributes for Transferable Assessment in Video
by: Ashutosh, Kumar, et al.
Published: (2025) -
ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos
by: Somayazulu, Arjun, et al.
Published: (2026) -
Mamba-VMR: Multimodal Query Augmentation via Generated Videos for Precise Temporal Grounding
by: Sun, Yunzhuo, et al.
Published: (2026)