Debiasing Central Fixation Confounds Reveals a Peripheral "Sweet Spot" for Human-like Scanpaths in Hard-Attention Vision
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Pengcheng, Shogo, Yonekura, Kuniyosh, Yasuo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Emergence of Fixational and Saccadic Movements in a Multi-Level Recurrent Attention Model for Vision
by: Pan, Pengcheng, et al.
Published: (2025)
by: Pan, Pengcheng, et al.
Published: (2025)
EVA: Bridging Performance and Human Alignment in Hard-Attention Vision Models for Image Classification
by: Pan, Pengcheng, et al.
Published: (2026)
by: Pan, Pengcheng, et al.
Published: (2026)
Not Too Generative, Not Too Discriminative: The Human Alignment Sweet Spot
by: Ortega, Jorge Chang, et al.
Published: (2026)
by: Ortega, Jorge Chang, et al.
Published: (2026)
Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction
by: Cartella, Giuseppe, et al.
Published: (2025)
by: Cartella, Giuseppe, et al.
Published: (2025)
Fairness-aware Vision Transformer via Debiased Self-Attention
by: Qiang, Yao, et al.
Published: (2023)
by: Qiang, Yao, et al.
Published: (2023)
AtteConDA: Attention-Based Conflict Suppression in Multi-Condition Diffusion Models and Synthetic Data Augmentation
by: Noguchi, Shogo
Published: (2026)
by: Noguchi, Shogo
Published: (2026)
Unified Dynamic Scanpath Predictors Outperform Individually Trained Neural Models
by: Abawi, Fares, et al.
Published: (2024)
by: Abawi, Fares, et al.
Published: (2024)
Unifying Top-down and Bottom-up Scanpath Prediction Using Transformers
by: Yang, Zhibo, et al.
Published: (2023)
by: Yang, Zhibo, et al.
Published: (2023)
Interpretable Debiasing of Vision-Language Models for Social Fairness
by: An, Na Min, et al.
Published: (2026)
by: An, Na Min, et al.
Published: (2026)
A Robotics-Inspired Scanpath Model Reveals the Importance of Uncertainty and Semantic Object Cues for Gaze Guidance in Dynamic Scenes
by: Mengers, Vito, et al.
Published: (2024)
by: Mengers, Vito, et al.
Published: (2024)
AGCD-Net: Attention Guided Context Debiasing Network for Emotion Recognition
by: Devi, Varsha, et al.
Published: (2025)
by: Devi, Varsha, et al.
Published: (2025)
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
by: Saini, Harshvardhan, et al.
Published: (2026)
by: Saini, Harshvardhan, et al.
Published: (2026)
Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation
by: Zhang, Ziheng, et al.
Published: (2025)
by: Zhang, Ziheng, et al.
Published: (2025)
AdaTP: Attention-Debiased Token Pruning for Video Large Language Models
by: Sun, Fengyuan, et al.
Published: (2025)
by: Sun, Fengyuan, et al.
Published: (2025)
Confounder-Aware Medical Data Selection for Fine-Tuning Pretrained Vision Models
by: Ji, Anyang, et al.
Published: (2025)
by: Ji, Anyang, et al.
Published: (2025)
Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding
by: Murlidaran, Shravan, et al.
Published: (2026)
by: Murlidaran, Shravan, et al.
Published: (2026)
Integrating Object Interaction Self-Attention and GAN-Based Debiasing for Visual Question Answering
by: Li, Zhifei, et al.
Published: (2025)
by: Li, Zhifei, et al.
Published: (2025)
A Unified Debiasing Approach for Vision-Language Models across Modalities and Tasks
by: Jung, Hoin, et al.
Published: (2024)
by: Jung, Hoin, et al.
Published: (2024)
Fair Generation without Unfair Distortions: Debiasing Text-to-Image Generation with Entanglement-Free Attention
by: Park, Jeonghoon, et al.
Published: (2025)
by: Park, Jeonghoon, et al.
Published: (2025)
Cognitive Alignment At No Cost: Inducing Human Attention Biases For Interpretable Vision Transformers
by: Knights, Ethan
Published: (2026)
by: Knights, Ethan
Published: (2026)
AmPLe: Supporting Vision-Language Models via Adaptive-Debiased Ensemble Multi-Prompt Learning
by: Song, Fei, et al.
Published: (2025)
by: Song, Fei, et al.
Published: (2025)
SPHERE: Unveiling Spatial Blind Spots in Vision-Language Models Through Hierarchical Evaluation
by: Zhang, Wenyu, et al.
Published: (2024)
by: Zhang, Wenyu, et al.
Published: (2024)
Human-like Object Grouping in Self-supervised Vision Transformers
by: Adeli, Hossein, et al.
Published: (2026)
by: Adeli, Hossein, et al.
Published: (2026)
Can Machines Imitate Humans? Integrative Turing-like tests for Language and Vision Demonstrate a Narrowing Gap
by: Zhang, Mengmi, et al.
Published: (2022)
by: Zhang, Mengmi, et al.
Published: (2022)
Doubly Debiased Test-Time Prompt Tuning for Vision-Language Models
by: Song, Fei, et al.
Published: (2025)
by: Song, Fei, et al.
Published: (2025)
TimeSearch: Hierarchical Video Search with Spotlight and Reflection for Human-like Long Video Understanding
by: Pan, Junwen, et al.
Published: (2025)
by: Pan, Junwen, et al.
Published: (2025)
Visual Data Diagnosis and Debiasing with Concept Graphs
by: Chakraborty, Rwiddhi, et al.
Published: (2024)
by: Chakraborty, Rwiddhi, et al.
Published: (2024)
Visible Yet Unreadable: A Systematic Blind Spot of Vision Language Models Across Writing Systems
by: Zhang, Jie, et al.
Published: (2025)
by: Zhang, Jie, et al.
Published: (2025)
Two-stage Vision Transformers and Hard Masking offer Robust Object Representations
by: Aniraj, Ananthu, et al.
Published: (2025)
by: Aniraj, Ananthu, et al.
Published: (2025)
The Linear Attention Resurrection in Vision Transformer
by: Zheng, Chuanyang
Published: (2025)
by: Zheng, Chuanyang
Published: (2025)
Improved Belief-Attention in Vision Task
by: Zhang, Guoqiang
Published: (2026)
by: Zhang, Guoqiang
Published: (2026)
Spiking Vision Transformer with Saccadic Attention
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
Privacy-Preserving Debiasing using Data Augmentation and Machine Unlearning
by: Pan, Zhixin, et al.
Published: (2024)
by: Pan, Zhixin, et al.
Published: (2024)
SEM: Sparse Embedding Modulation for Post-Hoc Debiasing of Vision-Language Models
by: Guimard, Quentin, et al.
Published: (2026)
by: Guimard, Quentin, et al.
Published: (2026)
When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs
by: Cao, Fanpu, et al.
Published: (2026)
by: Cao, Fanpu, et al.
Published: (2026)
Spot Risks Before Speaking! Unraveling Safety Attention Heads in Large Vision-Language Models
by: Zheng, Ziwei, et al.
Published: (2025)
by: Zheng, Ziwei, et al.
Published: (2025)
CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
by: Böhle, Moritz, et al.
Published: (2025)
by: Böhle, Moritz, et al.
Published: (2025)
Attention Guided CAM: Visual Explanations of Vision Transformer Guided by Self-Attention
by: Leem, Saebom, et al.
Published: (2024)
by: Leem, Saebom, et al.
Published: (2024)
Revealing Multi-View Hallucination in Large Vision-Language Models
by: Park, Wooje, et al.
Published: (2026)
by: Park, Wooje, et al.
Published: (2026)
Freeze and Reveal: Exposing Modality Bias in Vision-Language Models
by: Kavuri, Vivek Hruday, et al.
Published: (2025)
by: Kavuri, Vivek Hruday, et al.
Published: (2025)
Similar Items
-
Emergence of Fixational and Saccadic Movements in a Multi-Level Recurrent Attention Model for Vision
by: Pan, Pengcheng, et al.
Published: (2025) -
EVA: Bridging Performance and Human Alignment in Hard-Attention Vision Models for Image Classification
by: Pan, Pengcheng, et al.
Published: (2026) -
Not Too Generative, Not Too Discriminative: The Human Alignment Sweet Spot
by: Ortega, Jorge Chang, et al.
Published: (2026) -
Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction
by: Cartella, Giuseppe, et al.
Published: (2025) -
Fairness-aware Vision Transformer via Debiased Self-Attention
by: Qiang, Yao, et al.
Published: (2023)