HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Arif, Kazi Hasan Ibn, Yoon, JinYi, Nikolopoulos, Dimitrios S., Vandierendonck, Hans, John, Deepu, Ji, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference
by: Li, Xiangchen, et al.
Published: (2025)
by: Li, Xiangchen, et al.
Published: (2025)
iMAD: Intelligent Multi-Agent Debate for Efficient and Accurate LLM Inference
by: Fan, Wei, et al.
Published: (2025)
by: Fan, Wei, et al.
Published: (2025)
Polymorph: Energy-Efficient Multi-Label Classification for Video Streams on Embedded Devices
by: Ghafouri, Saeid, et al.
Published: (2025)
by: Ghafouri, Saeid, et al.
Published: (2025)
SLED: A Speculative LLM Decoding Framework for Efficient Edge Serving
by: Li, Xiangchen, et al.
Published: (2025)
by: Li, Xiangchen, et al.
Published: (2025)
PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model
by: Arif, Kazi Hasan Ibn, et al.
Published: (2025)
by: Arif, Kazi Hasan Ibn, et al.
Published: (2025)
TinyDrop: Tiny Model Guided Token Dropping for Vision Transformers
by: Wang, Guoxin, et al.
Published: (2025)
by: Wang, Guoxin, et al.
Published: (2025)
S2M3: Split-and-Share Multi-Modal Models for Distributed Multi-Task Inference on the Edge
by: Yoon, JinYi, et al.
Published: (2025)
by: Yoon, JinYi, et al.
Published: (2025)
LLM-Guided Runtime Parameter Optimization for Energy-Efficient Model Inference
by: Crumpacker, Katelyn, et al.
Published: (2026)
by: Crumpacker, Katelyn, et al.
Published: (2026)
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
by: Li, Xiangchen, et al.
Published: (2026)
by: Li, Xiangchen, et al.
Published: (2026)
MARVEL: An End-to-End Framework for Generating Model-Class Aware Custom RISC-V Extensions for Lightweight AI
by: M, Ajay Kumar, et al.
Published: (2025)
by: M, Ajay Kumar, et al.
Published: (2025)
EyeCue: Driver Cognitive Distraction Detection via Gaze-Empowered Egocentric Video Understanding
by: Zhang, Lang, et al.
Published: (2026)
by: Zhang, Lang, et al.
Published: (2026)
P3SL: Personalized Privacy-Preserving Split Learning on Heterogeneous Edge Devices
by: Fan, Wei, et al.
Published: (2025)
by: Fan, Wei, et al.
Published: (2025)
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
by: Li, Xiangchen, et al.
Published: (2026)
by: Li, Xiangchen, et al.
Published: (2026)
HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models
by: Liu, Jizhihui, et al.
Published: (2025)
by: Liu, Jizhihui, et al.
Published: (2025)
Harnessing Nonlinear Dynamics for Time-Driven Berry Phase in Classical Systems
by: Mahmood, Kazi T., et al.
Published: (2024)
by: Mahmood, Kazi T., et al.
Published: (2024)
Topological Insights from State Manipulation in a Classical Elastic System
by: Mahmood, Kazi T., et al.
Published: (2024)
by: Mahmood, Kazi T., et al.
Published: (2024)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
by: Wang, Guangtao, et al.
Published: (2025)
by: Wang, Guangtao, et al.
Published: (2025)
Attention Is not Everything: Efficient Alternatives for Vision
by: Kazi, Nur Mohammad, et al.
Published: (2026)
by: Kazi, Nur Mohammad, et al.
Published: (2026)
HERO: Rethinking Visual Token Early Dropping in High-Resolution Large Vision-Language Models
by: Li, Xu, et al.
Published: (2025)
by: Li, Xu, et al.
Published: (2025)
Topological Vibration Analysis of Elastic Lattices via Bloch Sphere Mapping
by: Mahmood, Kazi Tahsin, et al.
Published: (2025)
by: Mahmood, Kazi Tahsin, et al.
Published: (2025)
Equitable Skin Disease Prediction Using Transfer Learning and Domain Adaptation
by: Dip, Sajib Acharjee, et al.
Published: (2024)
by: Dip, Sajib Acharjee, et al.
Published: (2024)
HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit
by: Wu, Hao, et al.
Published: (2026)
by: Wu, Hao, et al.
Published: (2026)
STAR: Stage-Wise Attention-Guided Token Reduction for Efficient Large Vision-Language Models Inference
by: Guo, Yichen, et al.
Published: (2025)
by: Guo, Yichen, et al.
Published: (2025)
DiffPro: Joint Timestep and Layer-Wise Precision Optimization for Efficient Diffusion Inference
by: Amin, Farhana, et al.
Published: (2025)
by: Amin, Farhana, et al.
Published: (2025)
Less is More: Faster Maximum Clique Search by Work-Avoidance
by: Vandierendonck, Hans
Published: (2025)
by: Vandierendonck, Hans
Published: (2025)
FACT: Compositional Kernel Synthesis with a Three-Stage Agentic Workflow
by: Heidari, Sina, et al.
Published: (2026)
by: Heidari, Sina, et al.
Published: (2026)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
by: Fan, Jiakun, et al.
Published: (2025)
by: Fan, Jiakun, et al.
Published: (2025)
Modality Inflation: Energy Characterization and Optimization Opportunities for MLLM Inference
by: Moghadampanah, Mona, et al.
Published: (2025)
by: Moghadampanah, Mona, et al.
Published: (2025)
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
by: Liu, Ting, et al.
Published: (2024)
by: Liu, Ting, et al.
Published: (2024)
SecDTD: Dynamic Token Drop for Secure Transformers Inference
by: Cai, Yifei, et al.
Published: (2026)
by: Cai, Yifei, et al.
Published: (2026)
PADRe: A Unifying Polynomial Attention Drop-in Replacement for Efficient Vision Transformer
by: Letourneau, Pierre-David, et al.
Published: (2024)
by: Letourneau, Pierre-David, et al.
Published: (2024)
HEART-VIT: Hessian-Guided Efficient Dynamic Attention and Token Pruning in Vision Transformer
by: Uddin, Mohammad Helal, et al.
Published: (2025)
by: Uddin, Mohammad Helal, et al.
Published: (2025)
ProVerif Formal Verification Model for SSI-Mediated Wearable Health Data Sharing Architecture
by: Kazi Md Arif Shahriar
Published: (2026)
by: Kazi Md Arif Shahriar
Published: (2026)
Variation-aware Vision Token Dropping for Faster Large Vision-Language Models
by: Chen, Junjie, et al.
Published: (2025)
by: Chen, Junjie, et al.
Published: (2025)
Hi-Mamba: Hierarchical Mamba for Efficient Image Super-Resolution
by: Qiao, Junbo, et al.
Published: (2024)
by: Qiao, Junbo, et al.
Published: (2024)
Pyramid Token Pruning for High-Resolution Large Vision-Language Models via Region, Token, and Instruction-Guided Importance
by: Liang, Yuxuan, et al.
Published: (2025)
by: Liang, Yuxuan, et al.
Published: (2025)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
by: Jo, Dongwon, et al.
Published: (2026)
by: Jo, Dongwon, et al.
Published: (2026)
Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens
by: Shen, Meng, et al.
Published: (2026)
by: Shen, Meng, et al.
Published: (2026)
HiMat: DiT-based Ultra-High Resolution SVBRDF Generation
by: Wang, Zixiong, et al.
Published: (2025)
by: Wang, Zixiong, et al.
Published: (2025)
CATP: Cross-Attention Token Pruning for Accuracy Preserved Multimodal Model Inference
by: Liao, Ruqi, et al.
Published: (2024)
by: Liao, Ruqi, et al.
Published: (2024)
Similar Items
-
QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference
by: Li, Xiangchen, et al.
Published: (2025) -
iMAD: Intelligent Multi-Agent Debate for Efficient and Accurate LLM Inference
by: Fan, Wei, et al.
Published: (2025) -
Polymorph: Energy-Efficient Multi-Label Classification for Video Streams on Embedded Devices
by: Ghafouri, Saeid, et al.
Published: (2025) -
SLED: A Speculative LLM Decoding Framework for Efficient Edge Serving
by: Li, Xiangchen, et al.
Published: (2025) -
PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model
by: Arif, Kazi Hasan Ibn, et al.
Published: (2025)