Efficient Video Sampling: Pruning Temporally Redundant Tokens for Faster VLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Bagrov, Natan, Khvedchenia, Eugene, Tymchenko, Borys, Aharon, Shay, Kadoch, Lior, Keren, Tomer, Masad, Ofri, Geifman, Yonatan, Zilberstein, Ran, Rintamaki, Tuomas, Le, Matthieu, Tao, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Similarity-Aware Token Pruning: Your VLM but Faster
by: Jeddi, Ahmadreza, et al.
Published: (2025)
by: Jeddi, Ahmadreza, et al.
Published: (2025)
VGGHeads: 3D Multi Head Alignment with a Large-Scale Synthetic Dataset
by: Kupyn, Orest, et al.
Published: (2024)
by: Kupyn, Orest, et al.
Published: (2024)
Faster Cycle Detection in the Congested Clique
by: Censor-Hillel, Keren, et al.
Published: (2024)
by: Censor-Hillel, Keren, et al.
Published: (2024)
On a continuous Sárközy type problem
by: Kuca, Borys, et al.
Published: (2021)
by: Kuca, Borys, et al.
Published: (2021)
Faster Parameter-Efficient Tuning with Token Redundancy Reduction
by: Kim, Kwonyoung, et al.
Published: (2025)
by: Kim, Kwonyoung, et al.
Published: (2025)
SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding
by: Abramovich, Talor, et al.
Published: (2026)
by: Abramovich, Talor, et al.
Published: (2026)
Coherent Vorticity Extraction in 3D Homogeneous Isotropic Turbulence: Influence of the Reynolds Number and Geometrical Statistics
by: Benjamin Kadoch
Published: (2009)
by: Benjamin Kadoch
Published: (2009)
Think Clearly: Improving Reasoning via Redundant Token Pruning
by: Choi, Daewon, et al.
Published: (2025)
by: Choi, Daewon, et al.
Published: (2025)
VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
by: Wu, Zhenkai, et al.
Published: (2025)
by: Wu, Zhenkai, et al.
Published: (2025)
TokenTrim: Inference-Time Token Pruning for Autoregressive Long Video Generation
by: Shaulov, Ariel, et al.
Published: (2026)
by: Shaulov, Ariel, et al.
Published: (2026)
Chromatic Cardinalities via Redshift
by: Ben-Moshe, Shay, et al.
Published: (2023)
by: Ben-Moshe, Shay, et al.
Published: (2023)
Descent and cyclotomic redshift for chromatically localized algebraic K-theory
by: Ben-Moshe, Shay, et al.
Published: (2023)
by: Ben-Moshe, Shay, et al.
Published: (2023)
ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning
by: Lee, Yuna, et al.
Published: (2026)
by: Lee, Yuna, et al.
Published: (2026)
Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning
by: Huang, Yihong, et al.
Published: (2026)
by: Huang, Yihong, et al.
Published: (2026)
Yarn Ball Knots and Faster Computations
by: Bar-Natan, Dror, et al.
Published: (2021)
by: Bar-Natan, Dror, et al.
Published: (2021)
OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning
by: Li, Geng, et al.
Published: (2026)
by: Li, Geng, et al.
Published: (2026)
FFN Fusion: Rethinking Sequential Computation in Large Language Models
by: Bercovich, Akhiad, et al.
Published: (2025)
by: Bercovich, Akhiad, et al.
Published: (2025)
Is there Value in Reinforcement Learning?
by: Fox, Lior, et al.
Published: (2025)
by: Fox, Lior, et al.
Published: (2025)
Two pathways to resolve relational inconsistencies
by: Barak, Tomer, et al.
Published: (2024)
by: Barak, Tomer, et al.
Published: (2024)
Untrained neural networks can demonstrate memorization-independent abstract reasoning
by: Barak, Tomer, et al.
Published: (2024)
by: Barak, Tomer, et al.
Published: (2024)
AsymVLM: Asymmetric Token Pruning for Efficient Vision-Language Model Inference
by: Feng, Yilin, et al.
Published: (2026)
by: Feng, Yilin, et al.
Published: (2026)
COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
by: Kwek, Eugene, et al.
Published: (2025)
by: Kwek, Eugene, et al.
Published: (2025)
PACT: Pruning and Clustering-Based Token Reduction for Faster Visual Language Models
by: Dhouib, Mohamed, et al.
Published: (2025)
by: Dhouib, Mohamed, et al.
Published: (2025)
Coherent States of Systems with Quadratic Hamiltonians
by: V. G. Bagrov
Published: (2015)
by: V. G. Bagrov
Published: (2015)
Verification Required: The Impact of Information Credibility on AI Persuasion
by: Mahmud, Saaduddin, et al.
Published: (2026)
by: Mahmud, Saaduddin, et al.
Published: (2026)
Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity
by: Fang, Zhengyao, et al.
Published: (2026)
by: Fang, Zhengyao, et al.
Published: (2026)
Extending Puzzle for Mixture-of-Experts Reasoning Models with Application to GPT-OSS Acceleration
by: Bercovich, Akhiad, et al.
Published: (2026)
by: Bercovich, Akhiad, et al.
Published: (2026)
Experimental Realization of Rabi-Driven Reset for Fast Cooling of a High-Q Cavity
by: Blumenthal, Eliya, et al.
Published: (2026)
by: Blumenthal, Eliya, et al.
Published: (2026)
Fock State Generation and SWAP using a Rabi-Driven Qubit
by: Karaev, Natan, et al.
Published: (2026)
by: Karaev, Natan, et al.
Published: (2026)
SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
by: Ji, Yicheng, et al.
Published: (2025)
by: Ji, Yicheng, et al.
Published: (2025)
Faster Construction of a Planar Distance Oracle with Õ(1) Query Time
by: Boneh, Itai, et al.
Published: (2025)
by: Boneh, Itai, et al.
Published: (2025)
On the Trade-off between Redundancy and Local Coherence in Summarization
by: Cardenas, Ronald, et al.
Published: (2022)
by: Cardenas, Ronald, et al.
Published: (2022)
Faster or Stronger: Towards Flexible Visual Place Recognition via Weighted Aggregation and Token Pruning
by: Zeng, Zichao, et al.
Published: (2026)
by: Zeng, Zichao, et al.
Published: (2026)
Suppressing VLM Hallucinations with Spectral Representation Filtering
by: Ali, Ameen, et al.
Published: (2025)
by: Ali, Ameen, et al.
Published: (2025)
Higher Semiadditive Algebraic K-Theory and Redshift
by: Ben-Moshe, Shay, et al.
Published: (2021)
by: Ben-Moshe, Shay, et al.
Published: (2021)
Developing an Ontology for AI Act Fundamental Rights Impact Assessments
by: Rintamaki, Tytti, et al.
Published: (2024)
by: Rintamaki, Tytti, et al.
Published: (2024)
Towards An Automated AI Act FRIA Tool That Can Reuse GDPR's DPIA
by: Rintamaki, Tytti, et al.
Published: (2024)
by: Rintamaki, Tytti, et al.
Published: (2024)
REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space
by: Ashuach, Tomer, et al.
Published: (2024)
by: Ashuach, Tomer, et al.
Published: (2024)
On the Assessment of the Chilean Solar Thermal Regulation Using a Modular Simulation Model Coupled to a Multiobjective Optimization Algorithm
by: Jorge Contreras, et al.
Published: (2024)
by: Jorge Contreras, et al.
Published: (2024)
Faster Superword Tokenization
by: Schmidt, Craig W., et al.
Published: (2026)
by: Schmidt, Craig W., et al.
Published: (2026)
Similar Items
-
Similarity-Aware Token Pruning: Your VLM but Faster
by: Jeddi, Ahmadreza, et al.
Published: (2025) -
VGGHeads: 3D Multi Head Alignment with a Large-Scale Synthetic Dataset
by: Kupyn, Orest, et al.
Published: (2024) -
Faster Cycle Detection in the Congested Clique
by: Censor-Hillel, Keren, et al.
Published: (2024) -
On a continuous Sárközy type problem
by: Kuca, Borys, et al.
Published: (2021) -
Faster Parameter-Efficient Tuning with Token Redundancy Reduction
by: Kim, Kwonyoung, et al.
Published: (2025)