EVA: Bridging Performance and Human Alignment in Hard-Attention Vision Models for Image Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Pengcheng, Shogo, Yonekura, Yasuo, Kuniyoshi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Emergence of Fixational and Saccadic Movements in a Multi-Level Recurrent Attention Model for Vision
by: Pan, Pengcheng, et al.
Published: (2025)
by: Pan, Pengcheng, et al.
Published: (2025)
Debiasing Central Fixation Confounds Reveals a Peripheral "Sweet Spot" for Human-like Scanpaths in Hard-Attention Vision
by: Pan, Pengcheng, et al.
Published: (2026)
by: Pan, Pengcheng, et al.
Published: (2026)
Goal-conditioned dual-action imitation learning for dexterous dual-arm robot manipulation
by: Kim, Heecheol, et al.
Published: (2022)
by: Kim, Heecheol, et al.
Published: (2022)
Learning Conditionally Independent Transformations using Normal Subgroups in Group Theory
by: Nishitsunoi, Kayato, et al.
Published: (2025)
by: Nishitsunoi, Kayato, et al.
Published: (2025)
Evolution-aware VAriance (EVA) Coreset Selection for Medical Image Classification
by: Hong, Yuxin, et al.
Published: (2024)
by: Hong, Yuxin, et al.
Published: (2024)
Exploration-assisted Bottleneck Transition Toward Robust and Data-efficient Deformable Object Manipulation
by: Onishi, Yujiro, et al.
Published: (2026)
by: Onishi, Yujiro, et al.
Published: (2026)
Feature-Based Lie Group Transformer for Real-World Applications
by: Komatsu, Takayuki, et al.
Published: (2025)
by: Komatsu, Takayuki, et al.
Published: (2025)
Bridging Sensor Gaps via Attention Gated Tuning for Hyperspectral Image Classification
by: Xue, Xizhe, et al.
Published: (2023)
by: Xue, Xizhe, et al.
Published: (2023)
Enhancing Reusability of Learned Skills for Robot Manipulation via Gaze Information and Motion Bottlenecks
by: Takizawa, Ryo, et al.
Published: (2025)
by: Takizawa, Ryo, et al.
Published: (2025)
Lightweight Vision Transformer with Window and Spatial Attention for Food Image Classification
by: Gao, Xinle, et al.
Published: (2025)
by: Gao, Xinle, et al.
Published: (2025)
EVA: Mixture-of-Experts Semantic Variant Alignment for Compositional Zero-Shot Learning
by: Zhang, Xiao, et al.
Published: (2025)
by: Zhang, Xiao, et al.
Published: (2025)
AtteConDA: Attention-Based Conflict Suppression in Multi-Condition Diffusion Models and Synthetic Data Augmentation
by: Noguchi, Shogo
Published: (2026)
by: Noguchi, Shogo
Published: (2026)
A Saccade-inspired Approach to Image Classification using Vision Transformer Attention Maps
by: Dallain, Matthis, et al.
Published: (2026)
by: Dallain, Matthis, et al.
Published: (2026)
Attention Guided Alignment in Efficient Vision-Language Models
by: Mahajan, Shweta, et al.
Published: (2025)
by: Mahajan, Shweta, et al.
Published: (2025)
CARZero: Cross-Attention Alignment for Radiology Zero-Shot Classification
by: Lai, Haoran, et al.
Published: (2024)
by: Lai, Haoran, et al.
Published: (2024)
Hard Negative Sample Mining for Whole Slide Image Classification
by: Huang, Wentao, et al.
Published: (2024)
by: Huang, Wentao, et al.
Published: (2024)
Cognitive Alignment At No Cost: Inducing Human Attention Biases For Interpretable Vision Transformers
by: Knights, Ethan
Published: (2026)
by: Knights, Ethan
Published: (2026)
Soft-Hard Attention U-Net Model and Benchmark Dataset for Multiscale Image Shadow Removal
by: Cholopoulou, Eirini, et al.
Published: (2024)
by: Cholopoulou, Eirini, et al.
Published: (2024)
Exploring Interactive Semantic Alignment for Efficient HOI Detection with Vision-language Model
by: Dong, Jihao, et al.
Published: (2024)
by: Dong, Jihao, et al.
Published: (2024)
Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
by: Zhao, Shihao, et al.
Published: (2024)
by: Zhao, Shihao, et al.
Published: (2024)
Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models
by: Wang, Xuanhan, et al.
Published: (2025)
by: Wang, Xuanhan, et al.
Published: (2025)
Learning to Look: Cognitive Attention Alignment with Vision-Language Models
by: Yang, Ryan L., et al.
Published: (2025)
by: Yang, Ryan L., et al.
Published: (2025)
Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning
by: Chang, Aofei, et al.
Published: (2025)
by: Chang, Aofei, et al.
Published: (2025)
Large Language Models Facilitate Vision Reflection in Image Classification
by: An, Guoyuan, et al.
Published: (2025)
by: An, Guoyuan, et al.
Published: (2025)
Cosine-Normalized Attention for Hyperspectral Image Classification
by: Ahmad, Muhammad, et al.
Published: (2026)
by: Ahmad, Muhammad, et al.
Published: (2026)
Hyperbolic Image-and-Pointcloud Contrastive Learning for 3D Classification
by: Hu, Naiwen, et al.
Published: (2024)
by: Hu, Naiwen, et al.
Published: (2024)
Object-Conditioned Energy-Based Attention Map Alignment in Text-to-Image Diffusion Models
by: Zhang, Yasi, et al.
Published: (2024)
by: Zhang, Yasi, et al.
Published: (2024)
Bridging the Divide: Reconsidering Softmax and Linear Attention
by: Han, Dongchen, et al.
Published: (2024)
by: Han, Dongchen, et al.
Published: (2024)
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models
by: Johnson, Emily, et al.
Published: (2025)
by: Johnson, Emily, et al.
Published: (2025)
Optimizing Vision-Language Consistency via Cross-Layer Regional Attention Alignment
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
Optimizing Visual Question Answering Models for Driving: Bridging the Gap Between Human and Machine Attention Patterns
by: Rekanar, Kaavya, et al.
Published: (2024)
by: Rekanar, Kaavya, et al.
Published: (2024)
Data Adaptive Traceback for Vision-Language Foundation Models in Image Classification
by: Peng, Wenshuo, et al.
Published: (2024)
by: Peng, Wenshuo, et al.
Published: (2024)
Fusion of Foundation and Vision Transformer Model Features for Dermatoscopic Image Classification
by: Mahbod, Amirreza, et al.
Published: (2025)
by: Mahbod, Amirreza, et al.
Published: (2025)
Leveraging Vision-Language Models for Improving Domain Generalization in Image Classification
by: Addepalli, Sravanti, et al.
Published: (2023)
by: Addepalli, Sravanti, et al.
Published: (2023)
Vision Transformer for Classification of Breast Ultrasound Images
by: Gheflati, Behnaz, et al.
Published: (2021)
by: Gheflati, Behnaz, et al.
Published: (2021)
Vision Mamba for Classification of Breast Ultrasound Images
by: Nasiri-Sarvi, Ali, et al.
Published: (2024)
by: Nasiri-Sarvi, Ali, et al.
Published: (2024)
Bridging Human Evaluation to Infrared and Visible Image Fusion
by: Liu, Jinyuan, et al.
Published: (2026)
by: Liu, Jinyuan, et al.
Published: (2026)
Do Vision Transformers See Like Humans? Evaluating their Perceptual Alignment
by: Hernández-Cámara, Pablo, et al.
Published: (2025)
by: Hernández-Cámara, Pablo, et al.
Published: (2025)
Bridged Semantic Alignment for Zero-shot 3D Medical Image Diagnosis
by: Lai, Haoran, et al.
Published: (2025)
by: Lai, Haoran, et al.
Published: (2025)
WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models
by: Sugiura, Issa, et al.
Published: (2025)
by: Sugiura, Issa, et al.
Published: (2025)
Similar Items
-
Emergence of Fixational and Saccadic Movements in a Multi-Level Recurrent Attention Model for Vision
by: Pan, Pengcheng, et al.
Published: (2025) -
Debiasing Central Fixation Confounds Reveals a Peripheral "Sweet Spot" for Human-like Scanpaths in Hard-Attention Vision
by: Pan, Pengcheng, et al.
Published: (2026) -
Goal-conditioned dual-action imitation learning for dexterous dual-arm robot manipulation
by: Kim, Heecheol, et al.
Published: (2022) -
Learning Conditionally Independent Transformations using Normal Subgroups in Group Theory
by: Nishitsunoi, Kayato, et al.
Published: (2025) -
Evolution-aware VAriance (EVA) Coreset Selection for Medical Image Classification
by: Hong, Yuxin, et al.
Published: (2024)