Conscious Gaze: Adaptive Attention Mechanisms for Hallucination Mitigation in Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Bu, Weijue, Yuan, Guan, Zhang, Guixian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning for Autonomous Driving
by: Wang, Zhaohui, et al.
Published: (2025)
by: Wang, Zhaohui, et al.
Published: (2025)
Mitigating Catastrophic Forgetting in the Incremental Learning of Medical Images
by: Yavari, Sara, et al.
Published: (2025)
by: Yavari, Sara, et al.
Published: (2025)
Adaptive Self-Training for Object Detection
by: Vandeghen, Renaud, et al.
Published: (2022)
by: Vandeghen, Renaud, et al.
Published: (2022)
VITA: Zero-Shot Value Functions via Test-Time Adaptation of Vision-Language Models
by: Ziakas, Christos, et al.
Published: (2025)
by: Ziakas, Christos, et al.
Published: (2025)
CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models
by: Liu, Zhi
Published: (2026)
by: Liu, Zhi
Published: (2026)
ESCAPE: Energy-based Selective Adaptive Correction for Out-of-distribution 3D Human Pose Estimation
by: Bidulka, Luke, et al.
Published: (2024)
by: Bidulka, Luke, et al.
Published: (2024)
An Active Inference Model of Covert and Overt Visual Attention
by: Mišić, Tin, et al.
Published: (2025)
by: Mišić, Tin, et al.
Published: (2025)
ReLKD: Inter-Class Relation Learning with Knowledge Distillation for Generalized Category Discovery
by: Zhou, Fang, et al.
Published: (2025)
by: Zhou, Fang, et al.
Published: (2025)
Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders
by: Eymaël, Alexandre, et al.
Published: (2024)
by: Eymaël, Alexandre, et al.
Published: (2024)
Data Organization Matters in Multimodal Instruction Tuning: A Controlled Study of Capability Trade-offs
by: Tang, Guowei
Published: (2026)
by: Tang, Guowei
Published: (2026)
RDPO: Real Data Preference Optimization for Physics Consistency Video Generation
by: Qian, Wenxu, et al.
Published: (2025)
by: Qian, Wenxu, et al.
Published: (2025)
Invariant Representation via Decoupling Style and Spurious Features from Images
by: Li, Ruimeng, et al.
Published: (2023)
by: Li, Ruimeng, et al.
Published: (2023)
Learning to Seek Evidence: A Verifiable Reasoning Agent with Causal Faithfulness Analysis
by: Huang, Yuhang, et al.
Published: (2025)
by: Huang, Yuhang, et al.
Published: (2025)
FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion
by: Caselles-Dupré, Hugo, et al.
Published: (2026)
by: Caselles-Dupré, Hugo, et al.
Published: (2026)
EvoPrune: Early-Stage Visual Token Pruning for Efficient MLLMs
by: Chen, Yuhao, et al.
Published: (2026)
by: Chen, Yuhao, et al.
Published: (2026)
Pointing-Guided Target Estimation via Transformer-Based Attention
by: Müller, Luca, et al.
Published: (2025)
by: Müller, Luca, et al.
Published: (2025)
Pointing-Based Object Recognition
by: Hajdúch, Lukáš, et al.
Published: (2026)
by: Hajdúch, Lukáš, et al.
Published: (2026)
Analyzing Quality, Bias, and Performance in Text-to-Image Generative Models
by: Masrourisaadat, Nila, et al.
Published: (2024)
by: Masrourisaadat, Nila, et al.
Published: (2024)
FlightScope: An Experimental Comparative Review of Aircraft Detection Algorithms in Satellite Imagery
by: Ghazouali, Safouane El, et al.
Published: (2024)
by: Ghazouali, Safouane El, et al.
Published: (2024)
ForAug: Recombining Foregrounds and Backgrounds to Improve Vision Transformer Training with Bias Mitigation
by: Nauen, Tobias Christian, et al.
Published: (2025)
by: Nauen, Tobias Christian, et al.
Published: (2025)
PhysVid: Physics Aware Local Conditioning for Generative Video Models
by: Pathak, Saurabh, et al.
Published: (2026)
by: Pathak, Saurabh, et al.
Published: (2026)
Neuromorphic Monocular Depth Estimation with Uncertainty Modeling
by: Bergkvist, Viktor, et al.
Published: (2026)
by: Bergkvist, Viktor, et al.
Published: (2026)
Adapting Multimodal Foundation Models for Few-Shot Learning: A Comprehensive Study on Contrastive Captioners
by: Narasinghe, N. K. B. M. P. K. B., et al.
Published: (2025)
by: Narasinghe, N. K. B. M. P. K. B., et al.
Published: (2025)
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
by: Wu, Jason, et al.
Published: (2026)
by: Wu, Jason, et al.
Published: (2026)
Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models
by: Seo, Huichan, et al.
Published: (2025)
by: Seo, Huichan, et al.
Published: (2025)
Selection, Not Fusion: Radar-Modulated State Space Models for Radar-Camera Depth Estimation
by: Hou, Zhangcheng, et al.
Published: (2026)
by: Hou, Zhangcheng, et al.
Published: (2026)
Efficient Attention: Attention with Linear Complexities
by: Shen, Zhuoran, et al.
Published: (2018)
by: Shen, Zhuoran, et al.
Published: (2018)
Closed-Loop Neural Activation Control in Vision-Language-Action Models
by: Babu, Abhijith, et al.
Published: (2026)
by: Babu, Abhijith, et al.
Published: (2026)
Contrastive pretraining for semantic segmentation is robust to noisy positive pairs
by: Gerard, Sebastian, et al.
Published: (2022)
by: Gerard, Sebastian, et al.
Published: (2022)
Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design
by: Alabdulmohsin, Ibrahim, et al.
Published: (2023)
by: Alabdulmohsin, Ibrahim, et al.
Published: (2023)
Human-Centric Anomaly Detection in Surveillance Videos Using YOLO-World and Spatio-Temporal Deep Learning
by: Naeen, Mohammad Ali Etemadi, et al.
Published: (2025)
by: Naeen, Mohammad Ali Etemadi, et al.
Published: (2025)
Perceptual Flow Network for Visually Grounded Reasoning
by: Li, Yangfu, et al.
Published: (2026)
by: Li, Yangfu, et al.
Published: (2026)
Part-Level 3D Gaussian Vehicle Generation with Joint and Hinge Axis Estimation
by: Qian, Shiyao, et al.
Published: (2026)
by: Qian, Shiyao, et al.
Published: (2026)
Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning
by: Dong, Mingkang, et al.
Published: (2026)
by: Dong, Mingkang, et al.
Published: (2026)
Smooth regularization for efficient video recognition
by: Goldman, Gil, et al.
Published: (2025)
by: Goldman, Gil, et al.
Published: (2025)
Training a Student Expert via Semi-Supervised Foundation Model Distillation
by: Taghavi, Pardis, et al.
Published: (2026)
by: Taghavi, Pardis, et al.
Published: (2026)
nuScenes Knowledge Graph -- A comprehensive semantic representation of traffic scenes for trajectory prediction
by: Mlodzian, Leon, et al.
Published: (2023)
by: Mlodzian, Leon, et al.
Published: (2023)
In Context Learning with Vision Transformers: Case Study
by: Zhao, Antony, et al.
Published: (2025)
by: Zhao, Antony, et al.
Published: (2025)
Circuit Mechanisms for Spatial Relation Generation in Diffusion Transformers
by: Wang, Binxu, et al.
Published: (2026)
by: Wang, Binxu, et al.
Published: (2026)
MM-Food-100K: A 100,000-Sample Multimodal Food Intelligence Dataset with Verifiable Provenance
by: Dong, Yi, et al.
Published: (2025)
by: Dong, Yi, et al.
Published: (2025)
Similar Items
-
CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning for Autonomous Driving
by: Wang, Zhaohui, et al.
Published: (2025) -
Mitigating Catastrophic Forgetting in the Incremental Learning of Medical Images
by: Yavari, Sara, et al.
Published: (2025) -
Adaptive Self-Training for Object Detection
by: Vandeghen, Renaud, et al.
Published: (2022) -
VITA: Zero-Shot Value Functions via Test-Time Adaptation of Vision-Language Models
by: Ziakas, Christos, et al.
Published: (2025) -
CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models
by: Liu, Zhi
Published: (2026)