Enhancing Layer Attention Efficiency through Pruning Redundant Retrievals
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Hanze, Du, Yaosong, Yao, Zhibo, Zeng, Mengyao, Ge, Xiuqi, Huang, Xiande |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models
by: Tang, Kai, et al.
Published: (2025)
by: Tang, Kai, et al.
Published: (2025)
MAP: Mitigating Hallucinations in Large Vision-Language Models with Map-Level Attention Processing
by: Li, Chenxi, et al.
Published: (2025)
by: Li, Chenxi, et al.
Published: (2025)
Cloud Optical Thickness Retrievals Using Angle Invariant Attention Based Deep Learning Models
by: Tushar, Zahid Hassan, et al.
Published: (2025)
by: Tushar, Zahid Hassan, et al.
Published: (2025)
MagicWand: A Universal Agent for Generation and Evaluation Aligned with User Preference
by: Xu, Zitong, et al.
Published: (2025)
by: Xu, Zitong, et al.
Published: (2025)
AlterMOMA: Fusion Redundancy Pruning for Camera-LiDAR Fusion Models with Alternative Modality Masking
by: Sun, Shiqi, et al.
Published: (2024)
by: Sun, Shiqi, et al.
Published: (2024)
PLPHP: Per-Layer Per-Head Vision Token Pruning for Efficient Large Vision-Language Models
by: Meng, Yu, et al.
Published: (2025)
by: Meng, Yu, et al.
Published: (2025)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
by: Ge, Yuyao, et al.
Published: (2025)
by: Ge, Yuyao, et al.
Published: (2025)
Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs
by: Zhang, Qizhe, et al.
Published: (2025)
by: Zhang, Qizhe, et al.
Published: (2025)
Enhanced Structured Lasso Pruning with Class-wise Information
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
ESP-PCT: Enhanced VR Semantic Performance through Efficient Compression of Temporal and Spatial Redundancies in Point Cloud Transformers
by: Mei, Luoyu, et al.
Published: (2024)
by: Mei, Luoyu, et al.
Published: (2024)
SNP: Structured Neuron-level Pruning to Preserve Attention Scores
by: Shim, Kyunghwan, et al.
Published: (2024)
by: Shim, Kyunghwan, et al.
Published: (2024)
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy
by: Yang, Te, et al.
Published: (2024)
by: Yang, Te, et al.
Published: (2024)
FoPru: Focal Pruning for Efficient Large Vision-Language Models
by: Jiang, Lei, et al.
Published: (2024)
by: Jiang, Lei, et al.
Published: (2024)
Which Layer Causes Distribution Deviation? Entropy-Guided Adaptive Pruning for Diffusion and Flow Models
by: Li, Changlin, et al.
Published: (2025)
by: Li, Changlin, et al.
Published: (2025)
Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs
by: Zhang, Xiaofeng, et al.
Published: (2024)
by: Zhang, Xiaofeng, et al.
Published: (2024)
RainFusion: Adaptive Video Generation Acceleration via Multi-Dimensional Visual Redundancy
by: Chen, Aiyue, et al.
Published: (2025)
by: Chen, Aiyue, et al.
Published: (2025)
CLASP: Class-Adaptive Layer Fusion and Dual-Stage Pruning for Multimodal Large Language Models
by: Dang, Yunkai, et al.
Published: (2026)
by: Dang, Yunkai, et al.
Published: (2026)
SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models
by: Fang, Hengyu, et al.
Published: (2025)
by: Fang, Hengyu, et al.
Published: (2025)
AdaTP: Attention-Debiased Token Pruning for Video Large Language Models
by: Sun, Fengyuan, et al.
Published: (2025)
by: Sun, Fengyuan, et al.
Published: (2025)
Multi-Dimensional Pruning: Joint Channel, Layer and Block Pruning with Latency Constraint
by: Sun, Xinglong, et al.
Published: (2024)
by: Sun, Xinglong, et al.
Published: (2024)
PruneHal: Reducing Hallucinations in Multi-modal Large Language Models through Adaptive KV Cache Pruning
by: Sun, Fengyuan, et al.
Published: (2025)
by: Sun, Fengyuan, et al.
Published: (2025)
STAR: Stage-Wise Attention-Guided Token Reduction for Efficient Large Vision-Language Models Inference
by: Guo, Yichen, et al.
Published: (2025)
by: Guo, Yichen, et al.
Published: (2025)
EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent
by: Li, Jiaao, et al.
Published: (2025)
by: Li, Jiaao, et al.
Published: (2025)
Exploiting Information Redundancy in Attention Maps for Extreme Quantization of Vision Transformers
by: Maisonnave, Lucas, et al.
Published: (2025)
by: Maisonnave, Lucas, et al.
Published: (2025)
Understanding Pruning Regimes in Vision-Language Models Through Domain-Aware Layer Selection
by: Khaki, Saeed, et al.
Published: (2026)
by: Khaki, Saeed, et al.
Published: (2026)
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
by: Zhang, Qizhe, et al.
Published: (2024)
by: Zhang, Qizhe, et al.
Published: (2024)
FMCE-Net++: Feature Map Convergence Evaluation and Training
by: Zhu, Zhibo, et al.
Published: (2025)
by: Zhu, Zhibo, et al.
Published: (2025)
Parameter-Free Fine-tuning via Redundancy Elimination for Vision Foundation Models
by: Long, Jiahuan, et al.
Published: (2025)
by: Long, Jiahuan, et al.
Published: (2025)
Automatic Channel Pruning for Multi-Head Attention
by: Lee, Eunho, et al.
Published: (2024)
by: Lee, Eunho, et al.
Published: (2024)
DANet: Enhancing Small Object Detection through an Efficient Deformable Attention Network
by: Mia, Md Sohag, et al.
Published: (2023)
by: Mia, Md Sohag, et al.
Published: (2023)
Data Pruning by Information Maximization
by: Tan, Haoru, et al.
Published: (2025)
by: Tan, Haoru, et al.
Published: (2025)
SwinECAT: A Transformer-based fundus disease classification model with Shifted Window Attention and Efficient Channel Attention
by: Gu, Peiran, et al.
Published: (2025)
by: Gu, Peiran, et al.
Published: (2025)
Generalizable Facial Expression Recognition
by: Zhang, Yuhang, et al.
Published: (2024)
by: Zhang, Yuhang, et al.
Published: (2024)
UNSEEN: Enhancing Dataset Pruning from a Generalization Perspective
by: Xu, Furui, et al.
Published: (2025)
by: Xu, Furui, et al.
Published: (2025)
Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models
by: Apedo, Yvon, et al.
Published: (2026)
by: Apedo, Yvon, et al.
Published: (2026)
Rethinking Multi-Condition DiTs: Eliminating Redundant Attention via Position-Alignment and Keyword-Scoping
by: Zhou, Chao, et al.
Published: (2026)
by: Zhou, Chao, et al.
Published: (2026)
Llama Nemoretriever Colembed: Top-Performing Text-Image Retrieval Model
by: Xu, Mengyao, et al.
Published: (2025)
by: Xu, Mengyao, et al.
Published: (2025)
EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
by: Feng, Wenfeng, et al.
Published: (2025)
by: Feng, Wenfeng, et al.
Published: (2025)
GPTSee: Enhancing Moment Retrieval and Highlight Detection via Description-Based Similarity Features
by: Sun, Yunzhuo, et al.
Published: (2024)
by: Sun, Yunzhuo, et al.
Published: (2024)
Unveiling Redundancy in Diffusion Transformers (DiTs): A Systematic Study
by: Sun, Xibo, et al.
Published: (2024)
by: Sun, Xibo, et al.
Published: (2024)
Similar Items
-
Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models
by: Tang, Kai, et al.
Published: (2025) -
MAP: Mitigating Hallucinations in Large Vision-Language Models with Map-Level Attention Processing
by: Li, Chenxi, et al.
Published: (2025) -
Cloud Optical Thickness Retrievals Using Angle Invariant Attention Based Deep Learning Models
by: Tushar, Zahid Hassan, et al.
Published: (2025) -
MagicWand: A Universal Agent for Generation and Evaluation Aligned with User Preference
by: Xu, Zitong, et al.
Published: (2025) -
AlterMOMA: Fusion Redundancy Pruning for Camera-LiDAR Fusion Models with Alternative Modality Masking
by: Sun, Shiqi, et al.
Published: (2024)