Token Merging via Spatiotemporal Information Mining for Surgical Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Xixi, Yang, Chen, Zhang, Dong, Dong, Pingcheng, Yang, Xin, Cheng, Kwang-Ting
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914061714718720
author Jiang, Xixi
Yang, Chen
Zhang, Dong
Dong, Pingcheng
Yang, Xin
Cheng, Kwang-Ting
author_facet Jiang, Xixi
Yang, Chen
Zhang, Dong
Dong, Pingcheng
Yang, Xin
Cheng, Kwang-Ting
contents Vision Transformer models have shown impressive effectiveness in the surgical video understanding tasks through long-range dependency modeling. However, current methods suffer from prohibitive computational costs due to processing massive spatiotemporal tokens across video frames. While prior work on token merging has advanced model efficiency, they fail to adequately consider the inherent spatiotemporal structure of video data and overlook the heterogeneous nature of information distribution, leading to suboptimal performance. In this paper, we propose a spatiotemporal information mining token merging (STIM-TM) method, representing the first dedicated approach for surgical video understanding. STIM-TM introduces a decoupled strategy that reduces token redundancy along temporal and spatial dimensions independently. Specifically, the temporal component merges spatially corresponding tokens from consecutive frames using saliency weighting, preserving critical sequential information and maintaining continuity. Meanwhile, the spatial component prioritizes merging static tokens through temporal stability analysis, protecting dynamic regions containing essential surgical information. Operating in a training-free manner, STIM-TM achieves significant efficiency gains with over $65\%$ GFLOPs reduction while preserving competitive accuracy across comprehensive surgical video tasks. Our method also supports efficient training of long-sequence surgical videos, addressing computational bottlenecks in surgical applications.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23672
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Token Merging via Spatiotemporal Information Mining for Surgical Video Understanding
Jiang, Xixi
Yang, Chen
Zhang, Dong
Dong, Pingcheng
Yang, Xin
Cheng, Kwang-Ting
Computer Vision and Pattern Recognition
Vision Transformer models have shown impressive effectiveness in the surgical video understanding tasks through long-range dependency modeling. However, current methods suffer from prohibitive computational costs due to processing massive spatiotemporal tokens across video frames. While prior work on token merging has advanced model efficiency, they fail to adequately consider the inherent spatiotemporal structure of video data and overlook the heterogeneous nature of information distribution, leading to suboptimal performance. In this paper, we propose a spatiotemporal information mining token merging (STIM-TM) method, representing the first dedicated approach for surgical video understanding. STIM-TM introduces a decoupled strategy that reduces token redundancy along temporal and spatial dimensions independently. Specifically, the temporal component merges spatially corresponding tokens from consecutive frames using saliency weighting, preserving critical sequential information and maintaining continuity. Meanwhile, the spatial component prioritizes merging static tokens through temporal stability analysis, protecting dynamic regions containing essential surgical information. Operating in a training-free manner, STIM-TM achieves significant efficiency gains with over $65\%$ GFLOPs reduction while preserving competitive accuracy across comprehensive surgical video tasks. Our method also supports efficient training of long-sequence surgical videos, addressing computational bottlenecks in surgical applications.
title Token Merging via Spatiotemporal Information Mining for Surgical Video Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.23672