Input-Adaptive Visual Preprocessing for Efficient Fast Vision-Language Model Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Cahyani, Putu Indah Githa, Suartana, Komang David Dananjaya, Yudistira, Novanto |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Adaptive Fusion of Multimodal Deep Networks for Human Action Recognition
by: Yudistira, Novanto
Published: (2025)
by: Yudistira, Novanto
Published: (2025)
Efficient Object Detection of Marine Debris using Pruned YOLO Model
by: Aryaza, Abi, et al.
Published: (2025)
by: Aryaza, Abi, et al.
Published: (2025)
Adaptive Conformal Prediction for Reliable and Explainable Medical Image Classification
by: Octadion, One, et al.
Published: (2026)
by: Octadion, One, et al.
Published: (2026)
IndoHerb: Indonesia Medicinal Plants Recognition using Transfer Learning and Deep Learning
by: Musyaffa, Muhammad Salman Ikrar, et al.
Published: (2023)
by: Musyaffa, Muhammad Salman Ikrar, et al.
Published: (2023)
Training-Free Disentangled Text-Guided Image Editing via Sparse Latent Constraints
by: Shabrina, Mutiara, et al.
Published: (2025)
by: Shabrina, Mutiara, et al.
Published: (2025)
MAMI: Multi-Attentional Mutual-Information for Long Sequence Neuron Captioning
by: Fauzulhaq, Alfirsa Damasyifa, et al.
Published: (2024)
by: Fauzulhaq, Alfirsa Damasyifa, et al.
Published: (2024)
Illusion-Aware Visual Preprocessing and Anti-Illusion Prompting for Classic Illusion Understanding in Vision-Language Models
by: Zha, Junli, et al.
Published: (2026)
by: Zha, Junli, et al.
Published: (2026)
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
by: Zhang, Yuan, et al.
Published: (2024)
by: Zhang, Yuan, et al.
Published: (2024)
Harnessing Input-Adaptive Inference for Efficient VLN
by: Kang, Dongwoo, et al.
Published: (2025)
by: Kang, Dongwoo, et al.
Published: (2025)
Fast Inference of Visual Autoregressive Model with Adjacency-Adaptive Dynamical Draft Trees
by: Lei, Haodong, et al.
Published: (2025)
by: Lei, Haodong, et al.
Published: (2025)
Hybrid of DiffStride and Spectral Pooling in Convolutional Neural Networks
by: Rafif, Sulthan, et al.
Published: (2024)
by: Rafif, Sulthan, et al.
Published: (2024)
SparseSwin: Swin Transformer with Sparse Transformer Block
by: Pinasthika, Krisna, et al.
Published: (2023)
by: Pinasthika, Krisna, et al.
Published: (2023)
VIP: Visual-guided Prompt Evolution for Efficient Dense Vision-Language Inference
by: Zhu, Hao, et al.
Published: (2026)
by: Zhu, Hao, et al.
Published: (2026)
Improving Robustness of Vision-Language-Action Models by Restoring Corrupted Visual Inputs
by: Orjuela, Daniel Yezid Guarnizo, et al.
Published: (2026)
by: Orjuela, Daniel Yezid Guarnizo, et al.
Published: (2026)
EdgeFM: Efficient Edge Inference for Vision-Language Models
by: Deng, Mengling, et al.
Published: (2026)
by: Deng, Mengling, et al.
Published: (2026)
Adaptive High-Frequency Preprocessing for Video Coding
by: Pang, Yingxue, et al.
Published: (2025)
by: Pang, Yingxue, et al.
Published: (2025)
EVLM: An Efficient Vision-Language Model for Visual Understanding
by: Chen, Kaibing, et al.
Published: (2024)
by: Chen, Kaibing, et al.
Published: (2024)
Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models
by: He, Jialuo, et al.
Published: (2026)
by: He, Jialuo, et al.
Published: (2026)
Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference
by: Wang, Siyuan, et al.
Published: (2024)
by: Wang, Siyuan, et al.
Published: (2024)
VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models
by: Balakrishnan, Ravikumar, et al.
Published: (2025)
by: Balakrishnan, Ravikumar, et al.
Published: (2025)
VISOR: Visual Input-based Steering for Output Redirection in Vision-Language Models
by: Phute, Mansi, et al.
Published: (2025)
by: Phute, Mansi, et al.
Published: (2025)
Multi-Cue Adaptive Visual Token Pruning for Large Vision-Language Models
by: Luan, Bozhi, et al.
Published: (2025)
by: Luan, Bozhi, et al.
Published: (2025)
Adaptively Bypassing Vision Transformer Blocks for Efficient Visual Tracking
by: Yang, Xiangyang, et al.
Published: (2024)
by: Yang, Xiangyang, et al.
Published: (2024)
AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition
by: Lin, Zichuan, et al.
Published: (2025)
by: Lin, Zichuan, et al.
Published: (2025)
Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts
by: Rang, Miao, et al.
Published: (2025)
by: Rang, Miao, et al.
Published: (2025)
Event-Priori-Based Vision-Language Model for Efficient Visual Understanding
by: Qin, Haotong, et al.
Published: (2025)
by: Qin, Haotong, et al.
Published: (2025)
Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via Adaptive Token Skipping
by: Zeng, Weili, et al.
Published: (2025)
by: Zeng, Weili, et al.
Published: (2025)
AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference with Dynamical Text Guidance
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
Adaptive-VoCo: Complexity-Aware Visual Token Compression for Vision-Language Models
by: Guo, Xiaoyang, et al.
Published: (2025)
by: Guo, Xiaoyang, et al.
Published: (2025)
LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models
by: Debnath, Soumyaratna, et al.
Published: (2026)
by: Debnath, Soumyaratna, et al.
Published: (2026)
TRIO: Token Reduction via Inference-Objective Guidance for Efficient Vision-Language Models
by: Zhang, Haokui, et al.
Published: (2026)
by: Zhang, Haokui, et al.
Published: (2026)
Triage: Hierarchical Visual Budgeting for Efficient Video Reasoning in Vision-Language Models
by: Wang, Anmin, et al.
Published: (2026)
by: Wang, Anmin, et al.
Published: (2026)
Visual Prompt Engineering for Vision Language Models in Radiology
by: Denner, Stefan, et al.
Published: (2024)
by: Denner, Stefan, et al.
Published: (2024)
Towards Fast, Memory-based and Data-Efficient Vision-Language Policy
by: Li, Haoxuan, et al.
Published: (2025)
by: Li, Haoxuan, et al.
Published: (2025)
Saliency Driven Imagery Preprocessing for Efficient Compression -- Industrial Paper
by: Downes, Justin, et al.
Published: (2026)
by: Downes, Justin, et al.
Published: (2026)
Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models
by: Izzo, Riccardo Andrea, et al.
Published: (2026)
by: Izzo, Riccardo Andrea, et al.
Published: (2026)
Depth Adaptive Efficient Visual Autoregressive Modeling
by: Li, Chunliang, et al.
Published: (2026)
by: Li, Chunliang, et al.
Published: (2026)
A Preprocessing Framework for Video Machine Vision under Compression
by: Zhao, Fei, et al.
Published: (2025)
by: Zhao, Fei, et al.
Published: (2025)
ConsensusDrop: Fusing Visual and Cross-Modal Saliency for Efficient Vision Language Models
by: Parikh, Dhruv, et al.
Published: (2026)
by: Parikh, Dhruv, et al.
Published: (2026)
TurboVGGT: Fast Visual Geometry Reconstruction with Adaptive Alternating Attention
by: Huang, David, et al.
Published: (2026)
by: Huang, David, et al.
Published: (2026)
Similar Items
-
Towards Adaptive Fusion of Multimodal Deep Networks for Human Action Recognition
by: Yudistira, Novanto
Published: (2025) -
Efficient Object Detection of Marine Debris using Pruned YOLO Model
by: Aryaza, Abi, et al.
Published: (2025) -
Adaptive Conformal Prediction for Reliable and Explainable Medical Image Classification
by: Octadion, One, et al.
Published: (2026) -
IndoHerb: Indonesia Medicinal Plants Recognition using Transfer Learning and Deep Learning
by: Musyaffa, Muhammad Salman Ikrar, et al.
Published: (2023) -
Training-Free Disentangled Text-Guided Image Editing via Sparse Latent Constraints
by: Shabrina, Mutiara, et al.
Published: (2025)