Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Kumar, Yogesh |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
VideoLLM Benchmarks and Evaluation: A Survey
par: Kumar, Yogesh
Publié: (2025)
par: Kumar, Yogesh
Publié: (2025)
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
par: Qin, Jialong, et autres
Publié: (2025)
par: Qin, Jialong, et autres
Publié: (2025)
EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs
par: Li, Jiameng, et autres
Publié: (2026)
par: Li, Jiameng, et autres
Publié: (2026)
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
par: Wang, Han, et autres
Publié: (2024)
par: Wang, Han, et autres
Publié: (2024)
VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation
par: Wu, Shiwei, et autres
Publié: (2024)
par: Wu, Shiwei, et autres
Publié: (2024)
VideoLLM-online: Online Video Large Language Model for Streaming Video
par: Chen, Joya, et autres
Publié: (2024)
par: Chen, Joya, et autres
Publié: (2024)
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
par: Wang, Haibo, et autres
Publié: (2024)
par: Wang, Haibo, et autres
Publié: (2024)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
par: Li, Chenglin, et autres
Publié: (2026)
par: Li, Chenglin, et autres
Publié: (2026)
Lost in Time: A New Temporal Benchmark for VideoLLMs
par: Cores, Daniel, et autres
Publié: (2024)
par: Cores, Daniel, et autres
Publié: (2024)
Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
par: Cai, Jianfeng, et autres
Publié: (2025)
par: Cai, Jianfeng, et autres
Publié: (2025)
PruneVid: Visual Token Pruning for Efficient Video Large Language Models
par: Huang, Xiaohu, et autres
Publié: (2024)
par: Huang, Xiaohu, et autres
Publié: (2024)
Efficient Video Sampling: Pruning Temporally Redundant Tokens for Faster VLM Inference
par: Bagrov, Natan, et autres
Publié: (2025)
par: Bagrov, Natan, et autres
Publié: (2025)
Geometry-Guided Camera Motion Understanding in VideoLLMs
par: Feng, Haoan, et autres
Publié: (2026)
par: Feng, Haoan, et autres
Publié: (2026)
Geometry-Guided 3D Visual Token Pruning for Video-Language Models
par: Li, Han, et autres
Publié: (2026)
par: Li, Han, et autres
Publié: (2026)
Proact-VL: A Proactive VideoLLM for Real-Time AI Companions
par: Yan, Weicai, et autres
Publié: (2026)
par: Yan, Weicai, et autres
Publié: (2026)
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding
par: Fang, Pengcheng, et autres
Publié: (2025)
par: Fang, Pengcheng, et autres
Publié: (2025)
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
par: Guan, Yiran, et autres
Publié: (2026)
par: Guan, Yiran, et autres
Publié: (2026)
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
par: Liang, Yujia, et autres
Publié: (2025)
par: Liang, Yujia, et autres
Publié: (2025)
Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
par: Kim, Minji, et autres
Publié: (2025)
par: Kim, Minji, et autres
Publié: (2025)
HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models
par: Liu, Jizhihui, et autres
Publié: (2025)
par: Liu, Jizhihui, et autres
Publié: (2025)
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
par: Chatterjee, Dibyadip, et autres
Publié: (2025)
par: Chatterjee, Dibyadip, et autres
Publié: (2025)
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
par: Wang, Yueqian, et autres
Publié: (2024)
par: Wang, Yueqian, et autres
Publié: (2024)
EntropyPrune: Matrix Entropy Guided Visual Token Pruning for Multimodal Large Language Models
par: Wang, Yahong, et autres
Publié: (2026)
par: Wang, Yahong, et autres
Publié: (2026)
FastSTAR: Spatiotemporal Token Pruning for Efficient Autoregressive Video Synthesis
par: Yune, Sungwoong, et autres
Publié: (2026)
par: Yune, Sungwoong, et autres
Publié: (2026)
V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models
par: Lin, Xinying, et autres
Publié: (2026)
par: Lin, Xinying, et autres
Publié: (2026)
Video Patch Pruning: Efficient Video Instance Segmentation via Early Token Reduction
par: Glandorf, Patrick, et autres
Publié: (2026)
par: Glandorf, Patrick, et autres
Publié: (2026)
Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency
par: Li, Hongyu, et autres
Publié: (2025)
par: Li, Hongyu, et autres
Publié: (2025)
VLTP: Vision-Language Guided Token Pruning for Task-Oriented Segmentation
par: Chen, Hanning, et autres
Publié: (2024)
par: Chen, Hanning, et autres
Publié: (2024)
EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent
par: Li, Jiaao, et autres
Publié: (2025)
par: Li, Jiaao, et autres
Publié: (2025)
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
par: Kumar, Akash, et autres
Publié: (2025)
par: Kumar, Akash, et autres
Publié: (2025)
Temporal Object-Aware Vision Transformer for Few-Shot Video Object Detection
par: Kumar, Yogesh, et autres
Publié: (2025)
par: Kumar, Yogesh, et autres
Publié: (2025)
SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation
par: Ma, Shilin, et autres
Publié: (2026)
par: Ma, Shilin, et autres
Publié: (2026)
WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs
par: Zhang, Yulin, et autres
Publié: (2026)
par: Zhang, Yulin, et autres
Publié: (2026)
Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
par: Chung, Hyungjin, et autres
Publié: (2025)
par: Chung, Hyungjin, et autres
Publié: (2025)
TTF: Temporal Token Fusion for Efficient Video-Language Model
par: Huo, Simin, et autres
Publié: (2026)
par: Huo, Simin, et autres
Publié: (2026)
HEART-VIT: Hessian-Guided Efficient Dynamic Attention and Token Pruning in Vision Transformer
par: Uddin, Mohammad Helal, et autres
Publié: (2025)
par: Uddin, Mohammad Helal, et autres
Publié: (2025)
Keeping the Evidence Chain: Semantic Evidence Allocation for Training-Free Token Pruning in Video Temporal Grounding
par: Li, Jiaqi, et autres
Publié: (2026)
par: Li, Jiaqi, et autres
Publié: (2026)
Spatio-Temporal Token Pruning for Efficient High-Resolution GUI Agents
par: Xu, Zhou, et autres
Publié: (2026)
par: Xu, Zhou, et autres
Publié: (2026)
VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment
par: Kulkarni, Yogesh, et autres
Publié: (2025)
par: Kulkarni, Yogesh, et autres
Publié: (2025)
MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer
par: Cao, Jianjian, et autres
Publié: (2024)
par: Cao, Jianjian, et autres
Publié: (2024)
Documents similaires
-
VideoLLM Benchmarks and Evaluation: A Survey
par: Kumar, Yogesh
Publié: (2025) -
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
par: Qin, Jialong, et autres
Publié: (2025) -
EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs
par: Li, Jiameng, et autres
Publié: (2026) -
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
par: Wang, Han, et autres
Publié: (2024) -
VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation
par: Wu, Shiwei, et autres
Publié: (2024)