Do Vision Transformers See Like Humans? Evaluating their Perceptual Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Hernández-Cámara, Pablo, Jaén-Lorites, Jose Manuel, Vila-Tomás, Jorge, Laparra, Valero, Malo, Jesus |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Contrast Sensitivity in Multimodal Large Language Models: A Psychophysics-Inspired Evaluation
by: Hernández-Cámara, Pablo, et al.
Published: (2025)
by: Hernández-Cámara, Pablo, et al.
Published: (2025)
On the dynamic evolution of CLIP texture-shape bias and its relationship to human alignment and model robustness
by: Hernández-Cámara, Pablo, et al.
Published: (2025)
by: Hernández-Cámara, Pablo, et al.
Published: (2025)
Hues and Cues: Human vs. CLIP
by: Alabau-Bosque, Nuria, et al.
Published: (2025)
by: Alabau-Bosque, Nuria, et al.
Published: (2025)
From Images to Perception: Emergence of Perceptual Properties by Reconstructing Images
by: Hernández-Cámara, Pablo, et al.
Published: (2025)
by: Hernández-Cámara, Pablo, et al.
Published: (2025)
Parametric PerceptNet: A bio-inspired deep-net trained for Image Quality Assessment
by: Vila-Tomás, Jorge, et al.
Published: (2024)
by: Vila-Tomás, Jorge, et al.
Published: (2024)
Image Segmentation via Divisive Normalization: dealing with environmental diversity
by: Hernández-Cámara, Pablo, et al.
Published: (2024)
by: Hernández-Cámara, Pablo, et al.
Published: (2024)
Image Statistics Predict the Sensitivity of Perceptual Quality Metrics
by: Hepburn, Alexander, et al.
Published: (2023)
by: Hepburn, Alexander, et al.
Published: (2023)
A Turing Test for Artificial Nets devoted to model Human Vision
by: Vila-Tomás, Jorge, et al.
Published: (2025)
by: Vila-Tomás, Jorge, et al.
Published: (2025)
Parameter-Efficient Architectural Modifications for Translation-Invariant CNNs
by: Alabau-Bosque, Nuria, et al.
Published: (2026)
by: Alabau-Bosque, Nuria, et al.
Published: (2026)
Color Names in Vision-Language Models
by: Gomez-Villa, Alexandra, et al.
Published: (2025)
by: Gomez-Villa, Alexandra, et al.
Published: (2025)
Assessing invariance to affine transformations in image quality metrics
by: Alabau-Bosque, Nuria, et al.
Published: (2024)
by: Alabau-Bosque, Nuria, et al.
Published: (2024)
The Effect of Perceptual Metrics on Music Representation Learning for Genre Classification
by: Namgyal, Tashi, et al.
Published: (2024)
by: Namgyal, Tashi, et al.
Published: (2024)
Do MLLMs Exhibit Human-like Perceptual Behaviors? HVSBench: A Benchmark for MLLM Alignment with Human Perceptual Behavior
by: Lin, Jiaying, et al.
Published: (2024)
by: Lin, Jiaying, et al.
Published: (2024)
RAID-Database: human Responses to Affine Image Distortions
by: Daudén-Oliver, Paula, et al.
Published: (2024)
by: Daudén-Oliver, Paula, et al.
Published: (2024)
Learning to See Like Humans: Gaze-Aligned Cycling Safety Prediction
by: Perdigão, Luís Maria, et al.
Published: (2026)
by: Perdigão, Luís Maria, et al.
Published: (2026)
Do Vision Models Develop Human-Like Progressive Difficulty Understanding?
by: Huang, Zeyi, et al.
Published: (2025)
by: Huang, Zeyi, et al.
Published: (2025)
When Does Perceptual Alignment Benefit Vision Representations?
by: Sundaram, Shobhita, et al.
Published: (2024)
by: Sundaram, Shobhita, et al.
Published: (2024)
See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning
by: Zhang, Shuoshuo, et al.
Published: (2025)
by: Zhang, Shuoshuo, et al.
Published: (2025)
Beginning with You: Perceptual-Initialization Improves Vision-Language Representation and Alignment
by: Hu, Yang, et al.
Published: (2025)
by: Hu, Yang, et al.
Published: (2025)
VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression
by: Sargent, Kyle, et al.
Published: (2025)
by: Sargent, Kyle, et al.
Published: (2025)
Seeing is not Believing: An Identity Hider for Human Vision Privacy Protection
by: Wang, Tao, et al.
Published: (2023)
by: Wang, Tao, et al.
Published: (2023)
OmniEncoder: See, Hear, and Feel Continuous Motion Like Humans With One Encoder
by: Bai, Detao, et al.
Published: (2026)
by: Bai, Detao, et al.
Published: (2026)
Contour Integration Underlies Human-Like Vision
by: Lonnqvist, Ben, et al.
Published: (2025)
by: Lonnqvist, Ben, et al.
Published: (2025)
DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?
by: Zhou, Tianhong, et al.
Published: (2025)
by: Zhou, Tianhong, et al.
Published: (2025)
Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
by: Lee, Jonathan, et al.
Published: (2025)
by: Lee, Jonathan, et al.
Published: (2025)
What Do You See in Vehicle? Comprehensive Vision Solution for In-Vehicle Gaze Estimation
by: Cheng, Yihua, et al.
Published: (2024)
by: Cheng, Yihua, et al.
Published: (2024)
Do Vision-Language Models See Urban Scenes as People Do? An Urban Perception Benchmark
by: Mushkani, Rashid
Published: (2025)
by: Mushkani, Rashid
Published: (2025)
PP-Motion: Physical-Perceptual Fidelity Evaluation for Human Motion Generation
by: Zhao, Sihan, et al.
Published: (2025)
by: Zhao, Sihan, et al.
Published: (2025)
Achieving More Human Brain-Like Vision via Human EEG Representational Alignment
by: Lu, Zitong, et al.
Published: (2024)
by: Lu, Zitong, et al.
Published: (2024)
Vision-Language Models vs Human: Perceptual Image Quality Assessment
by: Mehmood, Imran, et al.
Published: (2026)
by: Mehmood, Imran, et al.
Published: (2026)
Do VLMs Need Vision Transformers? Evaluating State Space Models as Vision Encoders
by: Kuo, Shang-Jui Ray, et al.
Published: (2026)
by: Kuo, Shang-Jui Ray, et al.
Published: (2026)
Cognitive Alignment At No Cost: Inducing Human Attention Biases For Interpretable Vision Transformers
by: Knights, Ethan
Published: (2026)
by: Knights, Ethan
Published: (2026)
Alignment and Adversarial Robustness: Are More Human-Like Models More Secure?
by: Hoak, Blaine, et al.
Published: (2025)
by: Hoak, Blaine, et al.
Published: (2025)
See Fair, Speak Truth: Equitable Attention Improves Grounding and Reduces Hallucination in Vision-Language Alignment
by: Azeez, Mohammad Anas, et al.
Published: (2026)
by: Azeez, Mohammad Anas, et al.
Published: (2026)
Seeing Like Radiologists: Context- and Gaze-Guided Vision-Language Pretraining for Chest X-rays
by: Liu, Kang, et al.
Published: (2026)
by: Liu, Kang, et al.
Published: (2026)
Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
by: Shinnick, Zachary, et al.
Published: (2025)
by: Shinnick, Zachary, et al.
Published: (2025)
Do We Need Reformer for Vision? An Experimental Comparison with Vision Transformers
by: Bellaj, Ali El, et al.
Published: (2025)
by: Bellaj, Ali El, et al.
Published: (2025)
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
by: Yang, Fan, et al.
Published: (2026)
by: Yang, Fan, et al.
Published: (2026)
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
by: Wang, Beichen, et al.
Published: (2024)
by: Wang, Beichen, et al.
Published: (2024)
Do You See Me : A Multidimensional Benchmark for Evaluating Visual Perception in Multimodal LLMs
by: Kanade, Aditya, et al.
Published: (2025)
by: Kanade, Aditya, et al.
Published: (2025)
Similar Items
-
Contrast Sensitivity in Multimodal Large Language Models: A Psychophysics-Inspired Evaluation
by: Hernández-Cámara, Pablo, et al.
Published: (2025) -
On the dynamic evolution of CLIP texture-shape bias and its relationship to human alignment and model robustness
by: Hernández-Cámara, Pablo, et al.
Published: (2025) -
Hues and Cues: Human vs. CLIP
by: Alabau-Bosque, Nuria, et al.
Published: (2025) -
From Images to Perception: Emergence of Perceptual Properties by Reconstructing Images
by: Hernández-Cámara, Pablo, et al.
Published: (2025) -
Parametric PerceptNet: A bio-inspired deep-net trained for Image Quality Assessment
by: Vila-Tomás, Jorge, et al.
Published: (2024)