Learning to Look: Cognitive Attention Alignment with Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Ryan L., Bhusal, Dipkamal, Rastogi, Nidhi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do Sparse Subnetworks Exhibit Cognitively Aligned Attention? Effects of Pruning on Saliency Map Fidelity, Sparsity, and Concept Coherence
by: Suwal, Sanish, et al.
Published: (2025)
by: Suwal, Sanish, et al.
Published: (2025)
Concept-Based Masking: A Patch-Agnostic Defense Against Adversarial Patch Attacks
by: Mehrotra, Ayushi, et al.
Published: (2025)
by: Mehrotra, Ayushi, et al.
Published: (2025)
FACE: Faithful Automatic Concept Extraction
by: Bhusal, Dipkamal, et al.
Published: (2025)
by: Bhusal, Dipkamal, et al.
Published: (2025)
H-Sets: Hessian-Guided Discovery of Set-Level Feature Interactions in Image Classifiers
by: Mehrotra, Ayushi, et al.
Published: (2026)
by: Mehrotra, Ayushi, et al.
Published: (2026)
Training for Trustworthy Saliency Maps: Adversarial Training Meets Feature-Map Smoothing
by: Bhusal, Dipkamal, et al.
Published: (2026)
by: Bhusal, Dipkamal, et al.
Published: (2026)
Smaller is Better: Enhancing Transparency in Vehicle AI Systems via Pruning
by: Suwal, Sanish, et al.
Published: (2025)
by: Suwal, Sanish, et al.
Published: (2025)
Multi-Label Classification of Thoracic Diseases using Dense Convolutional Network on Chest Radiographs
by: Bhusal, Dipkamal, et al.
Published: (2022)
by: Bhusal, Dipkamal, et al.
Published: (2022)
Cognitive Alignment At No Cost: Inducing Human Attention Biases For Interpretable Vision Transformers
by: Knights, Ethan
Published: (2026)
by: Knights, Ethan
Published: (2026)
Safety Alignment for Vision Language Models
by: Liu, Zhendong, et al.
Published: (2024)
by: Liu, Zhendong, et al.
Published: (2024)
Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
by: Zhan, Yufei, et al.
Published: (2025)
by: Zhan, Yufei, et al.
Published: (2025)
Large Vision-Language Models Get Lost in Attention
by: Xi, Gongli, et al.
Published: (2026)
by: Xi, Gongli, et al.
Published: (2026)
Reasoning under Vision: Understanding Visual-Spatial Cognition in Vision-Language Models for CAPTCHA
by: Song, Python, et al.
Published: (2025)
by: Song, Python, et al.
Published: (2025)
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
by: Saini, Harshvardhan, et al.
Published: (2026)
by: Saini, Harshvardhan, et al.
Published: (2026)
Non-Contrastive Vision-Language Learning with Predictive Embedding Alignment
by: Kuhn, Lukas, et al.
Published: (2026)
by: Kuhn, Lukas, et al.
Published: (2026)
SPHINX: A Synthetic Environment for Visual Perception and Reasoning
by: Alam, Md Tanvirul, et al.
Published: (2025)
by: Alam, Md Tanvirul, et al.
Published: (2025)
Attention Prompting on Image for Large Vision-Language Models
by: Yu, Runpeng, et al.
Published: (2024)
by: Yu, Runpeng, et al.
Published: (2024)
Token-Level Inference-Time Alignment for Vision-Language Models
by: Chen, Kejia, et al.
Published: (2025)
by: Chen, Kejia, et al.
Published: (2025)
Subspace Alignment for Vision-Language Model Test-time Adaptation
by: Zeng, Zhichen, et al.
Published: (2026)
by: Zeng, Zhichen, et al.
Published: (2026)
Tuning Vision-Language Models with Candidate Labels by Prompt Alignment
by: Zhang, Zhifang, et al.
Published: (2024)
by: Zhang, Zhifang, et al.
Published: (2024)
Align Where the Words Look: Cross-Attention-Guided Patch Alignment with Contrastive and Transport Regularization for Bengali Captioning
by: Anonto, Riad Ahmed, et al.
Published: (2025)
by: Anonto, Riad Ahmed, et al.
Published: (2025)
Enhance Vision-Language Alignment with Noise
by: Huang, Sida, et al.
Published: (2024)
by: Huang, Sida, et al.
Published: (2024)
Batch Transformer: Look for Attention in Batch
by: Her, Myung Beom, et al.
Published: (2024)
by: Her, Myung Beom, et al.
Published: (2024)
Leveraging Vision-Language Models to Detect Attention in Educational Videos
by: Becquet, Gabriel, et al.
Published: (2026)
by: Becquet, Gabriel, et al.
Published: (2026)
A-VL: Adaptive Attention for Large Vision-Language Models
by: Zhang, Junyang, et al.
Published: (2024)
by: Zhang, Junyang, et al.
Published: (2024)
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
by: Woo, Sangmin, et al.
Published: (2024)
by: Woo, Sangmin, et al.
Published: (2024)
Enhancing Medical Large Vision-Language Models via Alignment Distillation
by: Chang, Aofei, et al.
Published: (2025)
by: Chang, Aofei, et al.
Published: (2025)
ITSELF: Attention Guided Fine-Grained Alignment for Vision-Language Retrieval
by: Nguyen, Tien-Huy, et al.
Published: (2026)
by: Nguyen, Tien-Huy, et al.
Published: (2026)
Evaluation of Safety Cognition Capability in Vision-Language Models for Autonomous Driving
by: Zhang, Enming, et al.
Published: (2025)
by: Zhang, Enming, et al.
Published: (2025)
FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language Alignment
by: Kim, Myunsoo, et al.
Published: (2025)
by: Kim, Myunsoo, et al.
Published: (2025)
Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models
by: Apedo, Yvon, et al.
Published: (2026)
by: Apedo, Yvon, et al.
Published: (2026)
Attention Hijacking: Response Manipulation Across Queries in Vision-Language Models
by: Wang, Zhiqiang, et al.
Published: (2026)
by: Wang, Zhiqiang, et al.
Published: (2026)
When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models
by: Yan, Yuping, et al.
Published: (2025)
by: Yan, Yuping, et al.
Published: (2025)
ATA: Bridging Implicit Reasoning with Attention-Guided and Action-Guided Inference for Vision-Language Action Models
by: Yang, Cheng, et al.
Published: (2026)
by: Yang, Cheng, et al.
Published: (2026)
Looking Back and Forth: Cross-Image Attention Calibration and Attentive Preference Learning for Multi-Image Hallucination Mitigation
by: Yang, Xiaochen, et al.
Published: (2026)
by: Yang, Xiaochen, et al.
Published: (2026)
SLIP: Structural-aware Language-Image Pretraining for Vision-Language Alignment
by: Lu, Wenbo
Published: (2025)
by: Lu, Wenbo
Published: (2025)
Mitigating Object Hallucinations in Large Vision-Language Models via Attention Calibration
by: Zhu, Younan, et al.
Published: (2025)
by: Zhu, Younan, et al.
Published: (2025)
All in One: Exploring Unified Vision-Language Tracking with Multi-Modal Alignment
by: Zhang, Chunhui, et al.
Published: (2023)
by: Zhang, Chunhui, et al.
Published: (2023)
CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
by: Böhle, Moritz, et al.
Published: (2025)
by: Böhle, Moritz, et al.
Published: (2025)
Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral Cloning
by: Qi, Xiuxiu, et al.
Published: (2025)
by: Qi, Xiuxiu, et al.
Published: (2025)
Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models
by: Yan, Hanqi, et al.
Published: (2025)
by: Yan, Hanqi, et al.
Published: (2025)
Similar Items
-
Do Sparse Subnetworks Exhibit Cognitively Aligned Attention? Effects of Pruning on Saliency Map Fidelity, Sparsity, and Concept Coherence
by: Suwal, Sanish, et al.
Published: (2025) -
Concept-Based Masking: A Patch-Agnostic Defense Against Adversarial Patch Attacks
by: Mehrotra, Ayushi, et al.
Published: (2025) -
FACE: Faithful Automatic Concept Extraction
by: Bhusal, Dipkamal, et al.
Published: (2025) -
H-Sets: Hessian-Guided Discovery of Set-Level Feature Interactions in Image Classifiers
by: Mehrotra, Ayushi, et al.
Published: (2026) -
Training for Trustworthy Saliency Maps: Adversarial Training Meets Feature-Map Smoothing
by: Bhusal, Dipkamal, et al.
Published: (2026)