HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: An, Joungbin, Grauman, Kristen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910086693126144
author An, Joungbin
Grauman, Kristen
author_facet An, Joungbin
Grauman, Kristen
contents Video temporal grounding, the task of localizing the start and end times of a natural language query in untrimmed video, requires capturing both global context and fine-grained temporal detail. This challenge is particularly pronounced in long videos, where existing methods often compromise temporal fidelity by over-downsampling or relying on fixed windows. We present HieraMamba, a hierarchical architecture that preserves temporal structure and semantic richness across scales. At its core are Anchor-MambaPooling (AMP) blocks, which utilize Mamba's selective scanning to produce compact anchor tokens that summarize video content at multiple granularities. Two complementary objectives, anchor-conditioned and segment-pooled contrastive losses, encourage anchors to retain local detail while remaining globally discriminative. HieraMamba sets a new state-of-the-art on Ego4D-NLQ, MAD, and TACoS, demonstrating precise, temporally faithful localization in long, untrimmed videos.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23043
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
An, Joungbin
Grauman, Kristen
Computer Vision and Pattern Recognition
Video temporal grounding, the task of localizing the start and end times of a natural language query in untrimmed video, requires capturing both global context and fine-grained temporal detail. This challenge is particularly pronounced in long videos, where existing methods often compromise temporal fidelity by over-downsampling or relying on fixed windows. We present HieraMamba, a hierarchical architecture that preserves temporal structure and semantic richness across scales. At its core are Anchor-MambaPooling (AMP) blocks, which utilize Mamba's selective scanning to produce compact anchor tokens that summarize video content at multiple granularities. Two complementary objectives, anchor-conditioned and segment-pooled contrastive losses, encourage anchors to retain local detail while remaining globally discriminative. HieraMamba sets a new state-of-the-art on Ego4D-NLQ, MAD, and TACoS, demonstrating precise, temporally faithful localization in long, untrimmed videos.
title HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.23043