Efficient Video Sampling: Pruning Temporally Redundant Tokens for Faster VLM Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bagrov, Natan, Khvedchenia, Eugene, Tymchenko, Borys, Aharon, Shay, Kadoch, Lior, Keren, Tomer, Masad, Ofri, Geifman, Yonatan, Zilberstein, Ran, Rintamaki, Tuomas, Le, Matthieu, Tao, Andrew
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908597662777344
author Bagrov, Natan
Khvedchenia, Eugene
Tymchenko, Borys
Aharon, Shay
Kadoch, Lior
Keren, Tomer
Masad, Ofri
Geifman, Yonatan
Zilberstein, Ran
Rintamaki, Tuomas
Le, Matthieu
Tao, Andrew
author_facet Bagrov, Natan
Khvedchenia, Eugene
Tymchenko, Borys
Aharon, Shay
Kadoch, Lior
Keren, Tomer
Masad, Ofri
Geifman, Yonatan
Zilberstein, Ran
Rintamaki, Tuomas
Le, Matthieu
Tao, Andrew
contents Vision-language models (VLMs) have recently expanded from static image understanding to video reasoning, but their scalability is fundamentally limited by the quadratic cost of processing dense frame sequences. Long videos often exceed the token budget of modern language models, leading to severe context limitations and latency issues. We introduce Efficient Video Sampling (EVS), a simple, plug-and-play method for reducing token redundancy in videos by identifying and pruning temporally static patches -- spatial regions that remain unchanged across consecutive frames. EVS preserves positional identity, requires no architectural changes or retraining. We show that EVS substantially reduces token count while maintaining semantic fidelity, enabling faster inference and longer input sequences. Applied at inference time, EVS reduces large language model (LLM) time-to-first-token (TTFT) by up to 4x with minimal accuracy loss. When combined with an uptraining phase using stochastic pruning rates, EVS yields models that are robust to varying compression levels and retain full performance under aggressive pruning. Extensive experiments demonstrate that EVS consistently improves efficiency-accuracy trade-offs, unlocking scalable video-language understanding without sacrificing quality.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14624
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Video Sampling: Pruning Temporally Redundant Tokens for Faster VLM Inference
Bagrov, Natan
Khvedchenia, Eugene
Tymchenko, Borys
Aharon, Shay
Kadoch, Lior
Keren, Tomer
Masad, Ofri
Geifman, Yonatan
Zilberstein, Ran
Rintamaki, Tuomas
Le, Matthieu
Tao, Andrew
Computer Vision and Pattern Recognition
Vision-language models (VLMs) have recently expanded from static image understanding to video reasoning, but their scalability is fundamentally limited by the quadratic cost of processing dense frame sequences. Long videos often exceed the token budget of modern language models, leading to severe context limitations and latency issues. We introduce Efficient Video Sampling (EVS), a simple, plug-and-play method for reducing token redundancy in videos by identifying and pruning temporally static patches -- spatial regions that remain unchanged across consecutive frames. EVS preserves positional identity, requires no architectural changes or retraining. We show that EVS substantially reduces token count while maintaining semantic fidelity, enabling faster inference and longer input sequences. Applied at inference time, EVS reduces large language model (LLM) time-to-first-token (TTFT) by up to 4x with minimal accuracy loss. When combined with an uptraining phase using stochastic pruning rates, EVS yields models that are robust to varying compression levels and retain full performance under aggressive pruning. Extensive experiments demonstrate that EVS consistently improves efficiency-accuracy trade-offs, unlocking scalable video-language understanding without sacrificing quality.
title Efficient Video Sampling: Pruning Temporally Redundant Tokens for Faster VLM Inference
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.14624