Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jinlong, Jiang, Liyuan, Zhang, Haonan, Sebe, Nicu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
by: Lin, Jinlong, et al.
Published: (2026)
by: Lin, Jinlong, et al.
Published: (2026)
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
by: Zhang, Pingping, et al.
Published: (2024)
by: Zhang, Pingping, et al.
Published: (2024)
Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding
by: Li, Jinlong, et al.
Published: (2025)
by: Li, Jinlong, et al.
Published: (2025)
Enhancing Robustness of Vision-Language Models through Orthogonality Learning and Self-Regularization
by: Li, Jinlong, et al.
Published: (2024)
by: Li, Jinlong, et al.
Published: (2024)
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
by: Gao, Shida, et al.
Published: (2025)
by: Gao, Shida, et al.
Published: (2025)
LESS: Label-Efficient and Single-Stage Referring 3D Segmentation
by: Liu, Xuexun, et al.
Published: (2024)
by: Liu, Xuexun, et al.
Published: (2024)
H$_{2}$OT: Hierarchical Hourglass Tokenizer for Efficient Video Pose Transformers
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
Democratizing Fine-grained Visual Recognition with Large Language Models
by: Liu, Mingxuan, et al.
Published: (2024)
by: Liu, Mingxuan, et al.
Published: (2024)
Large Language Models for Multimodal Deformable Image Registration
by: Ma, Mingrui, et al.
Published: (2024)
by: Ma, Mingrui, et al.
Published: (2024)
Orthogonal Projection Subspace to Aggregate Online Prior-knowledge for Continual Test-time Adaptation
by: Li, Jinlong, et al.
Published: (2025)
by: Li, Jinlong, et al.
Published: (2025)
FisherTune: Fisher-Guided Robust Tuning of Vision Foundation Models for Domain Generalized Segmentation
by: Zhao, Dong, et al.
Published: (2025)
by: Zhao, Dong, et al.
Published: (2025)
NullFace: Training-Free Localized Face Anonymization
by: Kung, Han-Wei, et al.
Published: (2025)
by: Kung, Han-Wei, et al.
Published: (2025)
Optimizing Resource Consumption in Diffusion Models through Hallucination Early Detection
by: Betti, Federico, et al.
Published: (2024)
by: Betti, Federico, et al.
Published: (2024)
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
by: Xing, Songlong, et al.
Published: (2025)
by: Xing, Songlong, et al.
Published: (2025)
Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism
by: Peruzzo, Elia, et al.
Published: (2025)
by: Peruzzo, Elia, et al.
Published: (2025)
3D Weakly Supervised Semantic Segmentation with 2D Vision-Language Guidance
by: Xu, Xiaoxu, et al.
Published: (2024)
by: Xu, Xiaoxu, et al.
Published: (2024)
Transferable-guided Attention Is All You Need for Video Domain Adaptation
by: Sacilotti, André, et al.
Published: (2024)
by: Sacilotti, André, et al.
Published: (2024)
Hourglass Tokenizer for Efficient Transformer-Based 3D Human Pose Estimation
by: Li, Wenhao, et al.
Published: (2023)
by: Li, Wenhao, et al.
Published: (2023)
TRIO: Token Reduction via Inference-Objective Guidance for Efficient Vision-Language Models
by: Zhang, Haokui, et al.
Published: (2026)
by: Zhang, Haokui, et al.
Published: (2026)
VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
by: Zhang, Ce, et al.
Published: (2025)
by: Zhang, Ce, et al.
Published: (2025)
Restore Anything Model via Efficient Degradation Adaptation
by: Ren, Bin, et al.
Published: (2024)
by: Ren, Bin, et al.
Published: (2024)
Multi-focal Conditioned Latent Diffusion for Person Image Synthesis
by: Liu, Jiaqi, et al.
Published: (2025)
by: Liu, Jiaqi, et al.
Published: (2025)
A Closer Look at Conditional Prompt Tuning for Vision-Language Models
by: Zhang, Ji, et al.
Published: (2025)
by: Zhang, Ji, et al.
Published: (2025)
Rethinking the Learning Paradigm for Facial Expression Recognition
by: Wang, Weijie, et al.
Published: (2022)
by: Wang, Weijie, et al.
Published: (2022)
Reverse Personalization
by: Kung, Han-Wei, et al.
Published: (2025)
by: Kung, Han-Wei, et al.
Published: (2025)
Towards End-to-End Explainable Facial Action Unit Recognition via Vision-Language Joint Learning
by: Ge, Xuri, et al.
Published: (2024)
by: Ge, Xuri, et al.
Published: (2024)
Efficient Multi-modal Large Language Models via Visual Token Grouping
by: Huang, Minbin, et al.
Published: (2024)
by: Huang, Minbin, et al.
Published: (2024)
Graph Transformer GANs with Graph Masked Modeling for Architectural Layout Generation
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
3D Weakly Supervised Semantic Segmentation via Class-Aware and Geometry-Guided Pseudo-Label Refinement
by: Xu, Xiaoxu, et al.
Published: (2025)
by: Xu, Xiaoxu, et al.
Published: (2025)
Video Patch Pruning: Efficient Video Instance Segmentation via Early Token Reduction
by: Glandorf, Patrick, et al.
Published: (2026)
by: Glandorf, Patrick, et al.
Published: (2026)
Safe Vision-Language Models via Unsafe Weights Manipulation
by: D'Incà, Moreno, et al.
Published: (2025)
by: D'Incà, Moreno, et al.
Published: (2025)
Large-scale Pre-trained Models are Surprisingly Strong in Incremental Novel Class Discovery
by: Liu, Mingxuan, et al.
Published: (2023)
by: Liu, Mingxuan, et al.
Published: (2023)
Hallucination Early Detection in Diffusion Models
by: Betti, Federico, et al.
Published: (2026)
by: Betti, Federico, et al.
Published: (2026)
Asymmetric GANs for Image-to-Image Translation
by: Tang, Hao, et al.
Published: (2019)
by: Tang, Hao, et al.
Published: (2019)
A Survey on Efficient Vision-Language-Action Models
by: Yu, Zhaoshu, et al.
Published: (2025)
by: Yu, Zhaoshu, et al.
Published: (2025)
RankFeat&RankWeight: Rank-1 Feature/Weight Removal for Out-of-distribution Detection
by: Song, Yue, et al.
Published: (2023)
by: Song, Yue, et al.
Published: (2023)
Uni4D: A Unified Self-Supervised Learning Framework for Point Cloud Videos
by: Zuo, Zhi, et al.
Published: (2025)
by: Zuo, Zhi, et al.
Published: (2025)
Vision+X: A Survey on Multimodal Learning in the Light of Data
by: Zhu, Ye, et al.
Published: (2022)
by: Zhu, Ye, et al.
Published: (2022)
Hierarchical Cross-Attention Network for Virtual Try-On
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
Similar Items
-
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
by: Lin, Jinlong, et al.
Published: (2026) -
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
by: Zhang, Pingping, et al.
Published: (2024) -
Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding
by: Li, Jinlong, et al.
Published: (2025) -
Enhancing Robustness of Vision-Language Models through Orthogonality Learning and Self-Regularization
by: Li, Jinlong, et al.
Published: (2024) -
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
by: Gao, Shida, et al.
Published: (2025)