Vision Transformer attention alignment with human visual perception in aesthetic object evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Carrasco, Miguel, González-Martín, César, Aranda, José, Oliveros, Luis |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Level of agreement between emotions generated by Artificial Intelligence and human evaluation: a methodological proposal
by: Carrasco, Miguel, et al.
Published: (2024)
by: Carrasco, Miguel, et al.
Published: (2024)
Self-supervised video pretraining yields robust and more human-aligned visual representations
by: Parthasarathy, Nikhil, et al.
Published: (2022)
by: Parthasarathy, Nikhil, et al.
Published: (2022)
Towards flexible perception with visual memory
by: Geirhos, Robert, et al.
Published: (2024)
by: Geirhos, Robert, et al.
Published: (2024)
CFM: Language-aligned Concept Foundation Model for Vision
by: Wittenmayer, Kai, et al.
Published: (2026)
by: Wittenmayer, Kai, et al.
Published: (2026)
VaPR -- Vision-language Preference alignment for Reasoning
by: Wadhawan, Rohan, et al.
Published: (2025)
by: Wadhawan, Rohan, et al.
Published: (2025)
Dimensions underlying the representational alignment of deep neural networks with humans
by: Mahner, Florian P., et al.
Published: (2024)
by: Mahner, Florian P., et al.
Published: (2024)
Evaluating alignment between humans and neural network representations in image-based learning tasks
by: Demircan, Can, et al.
Published: (2023)
by: Demircan, Can, et al.
Published: (2023)
Dilated Convolution with Learnable Spacings makes visual models more aligned with humans: a Grad-CAM study
by: Chamas, Rabih, et al.
Published: (2024)
by: Chamas, Rabih, et al.
Published: (2024)
QUEST: A robust attention formulation using query-modulated spherical attention
by: Govindarajan, Hariprasath, et al.
Published: (2026)
by: Govindarajan, Hariprasath, et al.
Published: (2026)
Moving object detection from multi-depth images with an attention-enhanced CNN
by: Shibukawa, Masato, et al.
Published: (2025)
by: Shibukawa, Masato, et al.
Published: (2025)
FasterViT: Fast Vision Transformers with Hierarchical Attention
by: Hatamizadeh, Ali, et al.
Published: (2023)
by: Hatamizadeh, Ali, et al.
Published: (2023)
From Ground to Air: Noise Robustness in Vision Transformers and CNNs for Event-Based Vehicle Classification with Potential UAV Applications
by: Almesafri, Nouf, et al.
Published: (2025)
by: Almesafri, Nouf, et al.
Published: (2025)
Improving Interpretation Faithfulness for Vision Transformers
by: Hu, Lijie, et al.
Published: (2023)
by: Hu, Lijie, et al.
Published: (2023)
Block-Recurrent Dynamics in Vision Transformers
by: Jacobs, Mozes, et al.
Published: (2025)
by: Jacobs, Mozes, et al.
Published: (2025)
Continual Adaptation of Vision Transformers for Federated Learning
by: Halbe, Shaunak, et al.
Published: (2023)
by: Halbe, Shaunak, et al.
Published: (2023)
Mechanisms of Non-Monotonic Scaling in Vision Transformers
by: Kumar, Anantha Padmanaban Krishna
Published: (2025)
by: Kumar, Anantha Padmanaban Krishna
Published: (2025)
Discovering Influential Neuron Path in Vision Transformers
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
DiffiT: Diffusion Vision Transformers for Image Generation
by: Hatamizadeh, Ali, et al.
Published: (2023)
by: Hatamizadeh, Ali, et al.
Published: (2023)
ADAPT to Robustify Prompt Tuning Vision Transformers
by: Eskandar, Masih, et al.
Published: (2024)
by: Eskandar, Masih, et al.
Published: (2024)
Accelerating Vision Transformers with Adaptive Patch Sizes
by: Choudhury, Rohan, et al.
Published: (2025)
by: Choudhury, Rohan, et al.
Published: (2025)
Class-Discriminative Attention Maps for Vision Transformers
by: Brocki, Lennart, et al.
Published: (2023)
by: Brocki, Lennart, et al.
Published: (2023)
Vision-language models lag human performance on physical dynamics and intent reasoning
by: Gu, Tianjun, et al.
Published: (2026)
by: Gu, Tianjun, et al.
Published: (2026)
MARS: Paying more attention to visual attributes for text-based person search
by: Ergasti, Alex, et al.
Published: (2024)
by: Ergasti, Alex, et al.
Published: (2024)
Intriguing Equivalence Structures of the Embedding Space of Vision Transformers
by: Salman, Shaeke, et al.
Published: (2024)
by: Salman, Shaeke, et al.
Published: (2024)
Oscillation-Reduced MXFP4 Training for Vision Transformers
by: Chen, Yuxiang, et al.
Published: (2025)
by: Chen, Yuxiang, et al.
Published: (2025)
Enhancing Vision Transformer Explainability Using Artificial Astrocytes
by: Echevarrieta-Catalan, Nicolas, et al.
Published: (2025)
by: Echevarrieta-Catalan, Nicolas, et al.
Published: (2025)
VariViT: A Vision Transformer for Variable Image Sizes
by: Varma, Aswathi, et al.
Published: (2026)
by: Varma, Aswathi, et al.
Published: (2026)
A Survey of the Self Supervised Learning Mechanisms for Vision Transformers
by: Khan, Asifullah, et al.
Published: (2024)
by: Khan, Asifullah, et al.
Published: (2024)
VisTabNet: Adapting Vision Transformers for Tabular Data
by: Wydmański, Witold, et al.
Published: (2024)
by: Wydmański, Witold, et al.
Published: (2024)
Beyond Scalars: Concept-Based Alignment Analysis in Vision Transformers
by: Vielhaben, Johanna, et al.
Published: (2024)
by: Vielhaben, Johanna, et al.
Published: (2024)
ScaleKD: Strong Vision Transformers Could Be Excellent Teachers
by: Fan, Jiawei, et al.
Published: (2024)
by: Fan, Jiawei, et al.
Published: (2024)
ScriptViT: Vision Transformer-Based Personalized Handwriting Generation
by: Acharya, Sajjan, et al.
Published: (2025)
by: Acharya, Sajjan, et al.
Published: (2025)
Pre-training Vision Transformers with Formula-driven Supervised Learning
by: Kataoka, Hirokatsu, et al.
Published: (2022)
by: Kataoka, Hirokatsu, et al.
Published: (2022)
Stratified Knowledge-Density Super-Network for Scalable Vision Transformers
by: Li, Longhua, et al.
Published: (2025)
by: Li, Longhua, et al.
Published: (2025)
What Helps---and What Hurts: Bidirectional Explanations for Vision Transformers
by: Su, Qin, et al.
Published: (2026)
by: Su, Qin, et al.
Published: (2026)
Octic Vision Transformers: Quicker ViTs Through Equivariance
by: Nordström, David, et al.
Published: (2025)
by: Nordström, David, et al.
Published: (2025)
Fairness-aware Vision Transformer via Debiased Self-Attention
by: Qiang, Yao, et al.
Published: (2023)
by: Qiang, Yao, et al.
Published: (2023)
PDiscoFormer: Relaxing Part Discovery Constraints with Vision Transformers
by: Aniraj, Ananthu, et al.
Published: (2024)
by: Aniraj, Ananthu, et al.
Published: (2024)
Efficient Few-Shot Learning in Remote Sensing: Fusing Vision and Vision-Language Models
by: Chua, Jia Yun, et al.
Published: (2025)
by: Chua, Jia Yun, et al.
Published: (2025)
Pig behavior dataset and Spatial-temporal perception and enhancement networks based on the attention mechanism for pig behavior recognition
by: Qi, Fangzheng, et al.
Published: (2025)
by: Qi, Fangzheng, et al.
Published: (2025)
Similar Items
-
Level of agreement between emotions generated by Artificial Intelligence and human evaluation: a methodological proposal
by: Carrasco, Miguel, et al.
Published: (2024) -
Self-supervised video pretraining yields robust and more human-aligned visual representations
by: Parthasarathy, Nikhil, et al.
Published: (2022) -
Towards flexible perception with visual memory
by: Geirhos, Robert, et al.
Published: (2024) -
CFM: Language-aligned Concept Foundation Model for Vision
by: Wittenmayer, Kai, et al.
Published: (2026) -
VaPR -- Vision-language Preference alignment for Reasoning
by: Wadhawan, Rohan, et al.
Published: (2025)