Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qin, Jialong, Zou, Xin, Lu, Di, Yan, Yibo, Hu, Xuming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
von: Kumar, Yogesh
Veröffentlicht: (2025)
von: Kumar, Yogesh
Veröffentlicht: (2025)
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs
von: Li, Jiameng, et al.
Veröffentlicht: (2026)
von: Li, Jiameng, et al.
Veröffentlicht: (2026)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
von: Chatterjee, Dibyadip, et al.
Veröffentlicht: (2025)
von: Chatterjee, Dibyadip, et al.
Veröffentlicht: (2025)
VideoLLM Benchmarks and Evaluation: A Survey
von: Kumar, Yogesh
Veröffentlicht: (2025)
von: Kumar, Yogesh
Veröffentlicht: (2025)
Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
von: Kim, Minji, et al.
Veröffentlicht: (2025)
von: Kim, Minji, et al.
Veröffentlicht: (2025)
Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
von: Cai, Jianfeng, et al.
Veröffentlicht: (2025)
von: Cai, Jianfeng, et al.
Veröffentlicht: (2025)
Geometry-Guided Camera Motion Understanding in VideoLLMs
von: Feng, Haoan, et al.
Veröffentlicht: (2026)
von: Feng, Haoan, et al.
Veröffentlicht: (2026)
VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation
von: Wu, Shiwei, et al.
Veröffentlicht: (2024)
von: Wu, Shiwei, et al.
Veröffentlicht: (2024)
WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs
von: Zhang, Yulin, et al.
Veröffentlicht: (2026)
von: Zhang, Yulin, et al.
Veröffentlicht: (2026)
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
von: Guan, Yiran, et al.
Veröffentlicht: (2026)
von: Guan, Yiran, et al.
Veröffentlicht: (2026)
Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
von: Chung, Hyungjin, et al.
Veröffentlicht: (2025)
von: Chung, Hyungjin, et al.
Veröffentlicht: (2025)
Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency
von: Li, Hongyu, et al.
Veröffentlicht: (2025)
von: Li, Hongyu, et al.
Veröffentlicht: (2025)
VideoLLM-online: Online Video Large Language Model for Streaming Video
von: Chen, Joya, et al.
Veröffentlicht: (2024)
von: Chen, Joya, et al.
Veröffentlicht: (2024)
Lost in Time: A New Temporal Benchmark for VideoLLMs
von: Cores, Daniel, et al.
Veröffentlicht: (2024)
von: Cores, Daniel, et al.
Veröffentlicht: (2024)
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding
von: Fang, Pengcheng, et al.
Veröffentlicht: (2025)
von: Fang, Pengcheng, et al.
Veröffentlicht: (2025)
Proact-VL: A Proactive VideoLLM for Real-Time AI Companions
von: Yan, Weicai, et al.
Veröffentlicht: (2026)
von: Yan, Weicai, et al.
Veröffentlicht: (2026)
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
von: Liang, Yujia, et al.
Veröffentlicht: (2025)
von: Liang, Yujia, et al.
Veröffentlicht: (2025)
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
"I Can See Forever!": Evaluating Real-time VideoLLMs for Assisting Individuals with Visual Impairments
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
DocPruner: A Storage-Efficient Framework for Multi-Vector Visual Document Retrieval via Adaptive Patch-Level Embedding Pruning
von: Yan, Yibo, et al.
Veröffentlicht: (2025)
von: Yan, Yibo, et al.
Veröffentlicht: (2025)
Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs
von: Kim, Kibum, et al.
Veröffentlicht: (2026)
von: Kim, Kibum, et al.
Veröffentlicht: (2026)
Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
von: Zou, Xin, et al.
Veröffentlicht: (2025)
von: Zou, Xin, et al.
Veröffentlicht: (2025)
Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
Less is More: Token-Efficient Video-QA via Adaptive Frame-Pruning and Semantic Graph Integration
von: Wang, Shaoguang, et al.
Veröffentlicht: (2025)
von: Wang, Shaoguang, et al.
Veröffentlicht: (2025)
EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent
von: Li, Jiaao, et al.
Veröffentlicht: (2025)
von: Li, Jiaao, et al.
Veröffentlicht: (2025)
PruneVid: Visual Token Pruning for Efficient Video Large Language Models
von: Huang, Xiaohu, et al.
Veröffentlicht: (2024)
von: Huang, Xiaohu, et al.
Veröffentlicht: (2024)
DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
Unified Spatiotemporal Token Compression for Video-LLMs at Ultra-Low Retention
von: Du, Junhao, et al.
Veröffentlicht: (2026)
von: Du, Junhao, et al.
Veröffentlicht: (2026)
Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding
von: Jin, Xinqi, et al.
Veröffentlicht: (2025)
von: Jin, Xinqi, et al.
Veröffentlicht: (2025)
GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning
von: Zhang, Jianghangfan, et al.
Veröffentlicht: (2025)
von: Zhang, Jianghangfan, et al.
Veröffentlicht: (2025)
Pruning the Unsurprising: Efficient LLM Reasoning via First-Token Surprisal
von: Zeng, Wenhao, et al.
Veröffentlicht: (2025)
von: Zeng, Wenhao, et al.
Veröffentlicht: (2025)
IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning
von: Sun, Zhichao, et al.
Veröffentlicht: (2026)
von: Sun, Zhichao, et al.
Veröffentlicht: (2026)
EvoPrune: Early-Stage Visual Token Pruning for Efficient MLLMs
von: Chen, Yuhao, et al.
Veröffentlicht: (2026)
von: Chen, Yuhao, et al.
Veröffentlicht: (2026)
Video Patch Pruning: Efficient Video Instance Segmentation via Early Token Reduction
von: Glandorf, Patrick, et al.
Veröffentlicht: (2026)
von: Glandorf, Patrick, et al.
Veröffentlicht: (2026)
Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models
von: Zheng, Kening, et al.
Veröffentlicht: (2024)
von: Zheng, Kening, et al.
Veröffentlicht: (2024)
SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
von: Li, Jungang, et al.
Veröffentlicht: (2024)
von: Li, Jungang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
von: Kumar, Yogesh
Veröffentlicht: (2025) -
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
von: Wang, Han, et al.
Veröffentlicht: (2024) -
EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs
von: Li, Jiameng, et al.
Veröffentlicht: (2026) -
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
von: Li, Chenglin, et al.
Veröffentlicht: (2026) -
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
von: Chatterjee, Dibyadip, et al.
Veröffentlicht: (2025)