Evaluating Graphical Perception Capabilities of Vision Transformers

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Poonam, Poonam, Vázquez, Pere-Pau, Ropinski, Timo
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915809525235712
author Poonam, Poonam
Vázquez, Pere-Pau
Ropinski, Timo
author_facet Poonam, Poonam
Vázquez, Pere-Pau
Ropinski, Timo
contents Vision Transformers, ViTs, have emerged as a powerful alternative to convolutional neural networks, CNNs, in a variety of image-based tasks. While CNNs have previously been evaluated for their ability to perform graphical perception tasks, which are essential for interpreting visualizations, the perceptual capabilities of ViTs remain largely unexplored. In this work, we investigate the performance of ViTs in elementary visual judgment tasks inspired by the foundational studies of Cleveland and McGill, which quantified the accuracy of human perception across different visual encodings. Inspired by their study, we benchmark ViTs against CNNs and human participants in a series of controlled graphical perception tasks. Our results reveal that, although ViTs demonstrate strong performance in general vision tasks, their alignment with human-like graphical perception in the visualization domain is limited. This study highlights key perceptual gaps and points to important considerations for the application of ViTs in visualization systems and graphical perceptual modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2602_18178
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Evaluating Graphical Perception Capabilities of Vision Transformers
Poonam, Poonam
Vázquez, Pere-Pau
Ropinski, Timo
Computer Vision and Pattern Recognition
Vision Transformers, ViTs, have emerged as a powerful alternative to convolutional neural networks, CNNs, in a variety of image-based tasks. While CNNs have previously been evaluated for their ability to perform graphical perception tasks, which are essential for interpreting visualizations, the perceptual capabilities of ViTs remain largely unexplored. In this work, we investigate the performance of ViTs in elementary visual judgment tasks inspired by the foundational studies of Cleveland and McGill, which quantified the accuracy of human perception across different visual encodings. Inspired by their study, we benchmark ViTs against CNNs and human participants in a series of controlled graphical perception tasks. Our results reveal that, although ViTs demonstrate strong performance in general vision tasks, their alignment with human-like graphical perception in the visualization domain is limited. This study highlights key perceptual gaps and points to important considerations for the application of ViTs in visualization systems and graphical perceptual modeling.
title Evaluating Graphical Perception Capabilities of Vision Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.18178