ApET: Approximation-Error Guided Token Compression for Efficient VLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Qiankun, Zhang, Ziyao, Wang, Haofei, Chen, Jie, Song, Zhen, Zheng, Hairong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Training-free Token Reduction for Vision Mamba
von: Ma, Qiankun, et al.
Veröffentlicht: (2025)
von: Ma, Qiankun, et al.
Veröffentlicht: (2025)
Towards Lossless Ultimate Vision Token Compression for VLMs
von: Zheng, Dehua, et al.
Veröffentlicht: (2025)
von: Zheng, Dehua, et al.
Veröffentlicht: (2025)
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
von: Wang, Ziyao, et al.
Veröffentlicht: (2026)
von: Wang, Ziyao, et al.
Veröffentlicht: (2026)
ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs
von: Gao, Yiling, et al.
Veröffentlicht: (2026)
von: Gao, Yiling, et al.
Veröffentlicht: (2026)
Beyond GSD-as-Token: Continuous Scale Conditioning for Remote Sensing VLMs
von: Zhang, Song, et al.
Veröffentlicht: (2026)
von: Zhang, Song, et al.
Veröffentlicht: (2026)
Contribution-aware Token Compression for Efficient Video Understanding via Reinforcement Learning
von: Ma, Yinchao, et al.
Veröffentlicht: (2026)
von: Ma, Yinchao, et al.
Veröffentlicht: (2026)
Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity
von: Fang, Zhengyao, et al.
Veröffentlicht: (2026)
von: Fang, Zhengyao, et al.
Veröffentlicht: (2026)
Balancing Saliency and Coverage: Semantic Prominence-Aware Budgeting for Visual Token Compression in VLMs
von: Lee, Jaehoon, et al.
Veröffentlicht: (2026)
von: Lee, Jaehoon, et al.
Veröffentlicht: (2026)
Beyond Shortcuts: Mitigating Visual Illusions in Frozen VLMs via Qualitative Reasoning
von: Guo, Hao, et al.
Veröffentlicht: (2026)
von: Guo, Hao, et al.
Veröffentlicht: (2026)
Language-Guided Token Compression with Reinforcement Learning in Large Vision-Language Models
von: Cao, Sihan, et al.
Veröffentlicht: (2026)
von: Cao, Sihan, et al.
Veröffentlicht: (2026)
AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
von: Zhang, Junyang, et al.
Veröffentlicht: (2025)
von: Zhang, Junyang, et al.
Veröffentlicht: (2025)
METok: Multi-Stage Event-based Token Compression for Efficient Long Video Understanding
von: Wang, Mengyue, et al.
Veröffentlicht: (2025)
von: Wang, Mengyue, et al.
Veröffentlicht: (2025)
QG-VTC: Question-Guided Visual Token Compression in MLLMs for Efficient VQA
von: Li, Shuai, et al.
Veröffentlicht: (2025)
von: Li, Shuai, et al.
Veröffentlicht: (2025)
Rethinking Token Reduction for Large Vision-Language Models
von: Wang, Yi, et al.
Veröffentlicht: (2026)
von: Wang, Yi, et al.
Veröffentlicht: (2026)
Motion Guided Token Compression for Efficient Masked Video Modeling
von: Feng, Yukun, et al.
Veröffentlicht: (2024)
von: Feng, Yukun, et al.
Veröffentlicht: (2024)
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models
von: Liu, Juntao, et al.
Veröffentlicht: (2025)
von: Liu, Juntao, et al.
Veröffentlicht: (2025)
TokenSeg: Efficient 3D Medical Image Segmentation via Hierarchical Visual Token Compression
von: Zeng, Sen, et al.
Veröffentlicht: (2026)
von: Zeng, Sen, et al.
Veröffentlicht: (2026)
CLiFT: Compressive Light-Field Tokens for Compute-Efficient and Adaptive Neural Rendering
von: Wang, Zhengqing, et al.
Veröffentlicht: (2025)
von: Wang, Zhengqing, et al.
Veröffentlicht: (2025)
QMoP: Query Guided Mixture-of-Projector for Efficient Visual Token Compression
von: Li, Zhongyang, et al.
Veröffentlicht: (2026)
von: Li, Zhongyang, et al.
Veröffentlicht: (2026)
Deep Pre-Alignment for VLMs
von: Yu, Tianyu, et al.
Veröffentlicht: (2026)
von: Yu, Tianyu, et al.
Veröffentlicht: (2026)
SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation
von: Ma, Shilin, et al.
Veröffentlicht: (2026)
von: Ma, Shilin, et al.
Veröffentlicht: (2026)
Token-level Correlation-guided Compression for Efficient Multimodal Document Understanding
von: Zhang, Renshan, et al.
Veröffentlicht: (2024)
von: Zhang, Renshan, et al.
Veröffentlicht: (2024)
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
von: Clark, Christopher, et al.
Veröffentlicht: (2026)
von: Clark, Christopher, et al.
Veröffentlicht: (2026)
Improving Autoregressive Image Generation through Coarse-to-Fine Token Prediction
von: Guo, Ziyao, et al.
Veröffentlicht: (2025)
von: Guo, Ziyao, et al.
Veröffentlicht: (2025)
CityLLaVA: Efficient Fine-Tuning for VLMs in City Scenario
von: Duan, Zhizhao, et al.
Veröffentlicht: (2024)
von: Duan, Zhizhao, et al.
Veröffentlicht: (2024)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
von: Zhang, Jianrui, et al.
Veröffentlicht: (2026)
von: Zhang, Jianrui, et al.
Veröffentlicht: (2026)
On the Concept Trustworthiness in Concept Bottleneck Models
von: Huang, Qihan, et al.
Veröffentlicht: (2024)
von: Huang, Qihan, et al.
Veröffentlicht: (2024)
Compressor-VLA: Instruction-Guided Visual Token Compression for Efficient Robotic Manipulation
von: Gao, Juntao, et al.
Veröffentlicht: (2025)
von: Gao, Juntao, et al.
Veröffentlicht: (2025)
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective
von: Lei, Lei, et al.
Veröffentlicht: (2025)
von: Lei, Lei, et al.
Veröffentlicht: (2025)
Recoverable Compression: A Multimodal Vision Token Recovery Mechanism Guided by Text Information
von: Chen, Yi, et al.
Veröffentlicht: (2024)
von: Chen, Yi, et al.
Veröffentlicht: (2024)
A Survey of Token Compression for Efficient Multimodal Large Language Models
von: Shao, Kele, et al.
Veröffentlicht: (2025)
von: Shao, Kele, et al.
Veröffentlicht: (2025)
LVC: A Lightweight Compression Framework for Enhancing VLMs in Long Video Understanding
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
MonoMVSNet: Monocular Priors Guided Multi-View Stereo Network
von: Jiang, Jianfei, et al.
Veröffentlicht: (2025)
von: Jiang, Jianfei, et al.
Veröffentlicht: (2025)
StreamingTOM: Streaming Token Compression for Efficient Video Understanding
von: Chen, Xueyi, et al.
Veröffentlicht: (2025)
von: Chen, Xueyi, et al.
Veröffentlicht: (2025)
TokenPacker: Efficient Visual Projector for Multimodal LLM
von: Li, Wentong, et al.
Veröffentlicht: (2024)
von: Li, Wentong, et al.
Veröffentlicht: (2024)
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning
von: li, Bonan, et al.
Veröffentlicht: (2025)
von: li, Bonan, et al.
Veröffentlicht: (2025)
RS3DBench: A Comprehensive Benchmark for 3D Spatial Perception in Remote Sensing
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
Switchable Token-Specific Codebook Quantization For Face Image Compression
von: Wang, Yongbo, et al.
Veröffentlicht: (2025)
von: Wang, Yongbo, et al.
Veröffentlicht: (2025)
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
von: Chinchure, Aditya, et al.
Veröffentlicht: (2025)
von: Chinchure, Aditya, et al.
Veröffentlicht: (2025)
Dynamic Token Compression for Efficient Video Understanding through Reinforcement Learning
von: Wang, Shida, et al.
Veröffentlicht: (2026)
von: Wang, Shida, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Training-free Token Reduction for Vision Mamba
von: Ma, Qiankun, et al.
Veröffentlicht: (2025) -
Towards Lossless Ultimate Vision Token Compression for VLMs
von: Zheng, Dehua, et al.
Veröffentlicht: (2025) -
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
von: Wang, Ziyao, et al.
Veröffentlicht: (2026) -
ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs
von: Gao, Yiling, et al.
Veröffentlicht: (2026) -
Beyond GSD-as-Token: Continuous Scale Conditioning for Remote Sensing VLMs
von: Zhang, Song, et al.
Veröffentlicht: (2026)