Vision Eagle Attention: a new lens for advancing image classification
Fuente:
arXiv
Salvato in:
| Autore principale: | Hasan, Mahmudul |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evolving CNN Architectures: From Custom Designs to Deep Residual Models for Diverse Image Classification and Detection Tasks
di: Hasan, Mahmudul, et al.
Pubblicazione: (2026)
di: Hasan, Mahmudul, et al.
Pubblicazione: (2026)
KAN-Mixers: a new deep learning architecture for image classification
di: Canuto, Jorge Luiz dos Santos, et al.
Pubblicazione: (2025)
di: Canuto, Jorge Luiz dos Santos, et al.
Pubblicazione: (2025)
MI-VisionShot: Few-shot adaptation of vision-language models for slide-level classification of histopathological images
di: Meseguer, Pablo, et al.
Pubblicazione: (2024)
di: Meseguer, Pablo, et al.
Pubblicazione: (2024)
Deep transfer learning for image classification: a survey
di: Plested, Jo, et al.
Pubblicazione: (2022)
di: Plested, Jo, et al.
Pubblicazione: (2022)
HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models
di: Arif, Kazi Hasan Ibn, et al.
Pubblicazione: (2024)
di: Arif, Kazi Hasan Ibn, et al.
Pubblicazione: (2024)
Identifying bias in CNN image classification using image scrambling and transforms
di: Erukude, Sai Teja
Pubblicazione: (2025)
di: Erukude, Sai Teja
Pubblicazione: (2025)
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
di: Li, Zhiqi, et al.
Pubblicazione: (2025)
di: Li, Zhiqi, et al.
Pubblicazione: (2025)
The Linear Attention Resurrection in Vision Transformer
di: Zheng, Chuanyang
Pubblicazione: (2025)
di: Zheng, Chuanyang
Pubblicazione: (2025)
Improved Belief-Attention in Vision Task
di: Zhang, Guoqiang
Pubblicazione: (2026)
di: Zhang, Guoqiang
Pubblicazione: (2026)
Spiking Vision Transformer with Saccadic Attention
di: Wang, Shuai, et al.
Pubblicazione: (2025)
di: Wang, Shuai, et al.
Pubblicazione: (2025)
Where are we with calibration under dataset shift in image classification?
di: Roschewitz, Mélanie, et al.
Pubblicazione: (2025)
di: Roschewitz, Mélanie, et al.
Pubblicazione: (2025)
SwinECAT: A Transformer-based fundus disease classification model with Shifted Window Attention and Efficient Channel Attention
di: Gu, Peiran, et al.
Pubblicazione: (2025)
di: Gu, Peiran, et al.
Pubblicazione: (2025)
Attention Guided CAM: Visual Explanations of Vision Transformer Guided by Self-Attention
di: Leem, Saebom, et al.
Pubblicazione: (2024)
di: Leem, Saebom, et al.
Pubblicazione: (2024)
CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
di: Böhle, Moritz, et al.
Pubblicazione: (2025)
di: Böhle, Moritz, et al.
Pubblicazione: (2025)
Revisiting the Integration of Convolution and Attention for Vision Backbone
di: Zhu, Lei, et al.
Pubblicazione: (2024)
di: Zhu, Lei, et al.
Pubblicazione: (2024)
Attention Retention for Continual Learning with Vision Transformers
di: Lu, Yue, et al.
Pubblicazione: (2026)
di: Lu, Yue, et al.
Pubblicazione: (2026)
FedDropoutAvg: Generalizable federated learning for histopathology image classification
di: Gunesli, Gozde N., et al.
Pubblicazione: (2021)
di: Gunesli, Gozde N., et al.
Pubblicazione: (2021)
Deep Learning for Breast Cancer Detection: Comparative Analysis of ConvNeXT and EfficientNet
di: Hasan, Mahmudul
Pubblicazione: (2025)
di: Hasan, Mahmudul
Pubblicazione: (2025)
Motion-enhanced Cardiac Anatomy Segmentation via an Insertable Temporal Attention Module
di: Hasan, Md. Kamrul, et al.
Pubblicazione: (2025)
di: Hasan, Md. Kamrul, et al.
Pubblicazione: (2025)
Efficient Event-Based Object Detection: A Hybrid Neural Network with Spatial and Temporal Attention
di: Ahmed, Soikat Hasan, et al.
Pubblicazione: (2024)
di: Ahmed, Soikat Hasan, et al.
Pubblicazione: (2024)
Attention mechanisms and transfer learning for robust peach leaf damage classification under domain shift
di: Cánovas-Rodriguez, Adrián, et al.
Pubblicazione: (2026)
di: Cánovas-Rodriguez, Adrián, et al.
Pubblicazione: (2026)
Mitigating annotation shift in cancer classification using single image generative models
di: Arcas, Marta Buetas, et al.
Pubblicazione: (2024)
di: Arcas, Marta Buetas, et al.
Pubblicazione: (2024)
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
di: Woo, Sangmin, et al.
Pubblicazione: (2024)
di: Woo, Sangmin, et al.
Pubblicazione: (2024)
Vision Augmentation Prediction Autoencoder with Attention Design (VAPAAD)
di: Yin, Yiqiao
Pubblicazione: (2024)
di: Yin, Yiqiao
Pubblicazione: (2024)
Attention Prompting on Image for Large Vision-Language Models
di: Yu, Runpeng, et al.
Pubblicazione: (2024)
di: Yu, Runpeng, et al.
Pubblicazione: (2024)
Large Vision-Language Models Get Lost in Attention
di: Xi, Gongli, et al.
Pubblicazione: (2026)
di: Xi, Gongli, et al.
Pubblicazione: (2026)
Rethinking Causal Mask Attention for Vision-Language Inference
di: Pei, Xiaohuan, et al.
Pubblicazione: (2025)
di: Pei, Xiaohuan, et al.
Pubblicazione: (2025)
Accelerating Local AI on Consumer GPUs: A Hardware-Aware Dynamic Strategy for YOLOv10s
di: Masum, Mahmudul Islam, et al.
Pubblicazione: (2025)
di: Masum, Mahmudul Islam, et al.
Pubblicazione: (2025)
Exploiting LMM-based knowledge for image classification tasks
di: Tzelepi, Maria, et al.
Pubblicazione: (2024)
di: Tzelepi, Maria, et al.
Pubblicazione: (2024)
Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models
di: Apedo, Yvon, et al.
Pubblicazione: (2026)
di: Apedo, Yvon, et al.
Pubblicazione: (2026)
Performance of computer vision algorithms for fine-grained classification using crowdsourced insect images
di: Pucci, Rita, et al.
Pubblicazione: (2024)
di: Pucci, Rita, et al.
Pubblicazione: (2024)
A-VL: Adaptive Attention for Large Vision-Language Models
di: Zhang, Junyang, et al.
Pubblicazione: (2024)
di: Zhang, Junyang, et al.
Pubblicazione: (2024)
SPARO: Selective Attention for Robust and Compositional Transformer Encodings for Vision
di: Vani, Ankit, et al.
Pubblicazione: (2024)
di: Vani, Ankit, et al.
Pubblicazione: (2024)
Leveraging Vision-Language Models to Detect Attention in Educational Videos
di: Becquet, Gabriel, et al.
Pubblicazione: (2026)
di: Becquet, Gabriel, et al.
Pubblicazione: (2026)
Learning to Look: Cognitive Attention Alignment with Vision-Language Models
di: Yang, Ryan L., et al.
Pubblicazione: (2025)
di: Yang, Ryan L., et al.
Pubblicazione: (2025)
PolaFormer: Polarity-aware Linear Attention for Vision Transformers
di: Meng, Weikang, et al.
Pubblicazione: (2025)
di: Meng, Weikang, et al.
Pubblicazione: (2025)
A review of deep learning-based information fusion techniques for multimodal medical image classification
di: Li, Yihao, et al.
Pubblicazione: (2024)
di: Li, Yihao, et al.
Pubblicazione: (2024)
Emergence of Fixational and Saccadic Movements in a Multi-Level Recurrent Attention Model for Vision
di: Pan, Pengcheng, et al.
Pubblicazione: (2025)
di: Pan, Pengcheng, et al.
Pubblicazione: (2025)
Human-annotated label noise and their impact on ConvNets for remote sensing image scene classification
di: Peng, Longkang, et al.
Pubblicazione: (2023)
di: Peng, Longkang, et al.
Pubblicazione: (2023)
Attention Hijacking: Response Manipulation Across Queries in Vision-Language Models
di: Wang, Zhiqiang, et al.
Pubblicazione: (2026)
di: Wang, Zhiqiang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Evolving CNN Architectures: From Custom Designs to Deep Residual Models for Diverse Image Classification and Detection Tasks
di: Hasan, Mahmudul, et al.
Pubblicazione: (2026) -
KAN-Mixers: a new deep learning architecture for image classification
di: Canuto, Jorge Luiz dos Santos, et al.
Pubblicazione: (2025) -
MI-VisionShot: Few-shot adaptation of vision-language models for slide-level classification of histopathological images
di: Meseguer, Pablo, et al.
Pubblicazione: (2024) -
Deep transfer learning for image classification: a survey
di: Plested, Jo, et al.
Pubblicazione: (2022) -
HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models
di: Arif, Kazi Hasan Ibn, et al.
Pubblicazione: (2024)