Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Kumar, Yogesh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908500868726784
author Kumar, Yogesh
author_facet Kumar, Yogesh
contents Vision Language Models (VLMs) struggle with long-form videos due to the quadratic complexity of attention mechanisms. We propose Language-Guided Temporal Token Pruning (LGTTP), which leverages temporal cues from queries to adaptively prune video tokens, preserving contextual continuity while reducing computational overhead. Unlike uniform pruning or keyframe selection, LGTTP retains higher token density in temporally relevant segments. Our model-agnostic framework integrates with TimeChat and LLaVA-Video, achieving a 65% reduction in computation while preserving 97-99% of the original performance. On QVHighlights, LGTTP improves HIT@1 by +9.5%, and on Charades-STA, it retains 99.6% of R@1. It excels on queries with explicit temporal markers and remains effective across general video understanding tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17686
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
Kumar, Yogesh
Computer Vision and Pattern Recognition
Vision Language Models (VLMs) struggle with long-form videos due to the quadratic complexity of attention mechanisms. We propose Language-Guided Temporal Token Pruning (LGTTP), which leverages temporal cues from queries to adaptively prune video tokens, preserving contextual continuity while reducing computational overhead. Unlike uniform pruning or keyframe selection, LGTTP retains higher token density in temporally relevant segments. Our model-agnostic framework integrates with TimeChat and LLaVA-Video, achieving a 65% reduction in computation while preserving 97-99% of the original performance. On QVHighlights, LGTTP improves HIT@1 by +9.5%, and on Charades-STA, it retains 99.6% of R@1. It excels on queries with explicit temporal markers and remains effective across general video understanding tasks.
title Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.17686