When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ortu, Francesco, Jin, Zhijing, Doimo, Diego, Cazzaniga, Alberto |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs
von: Khayatan, Pegah, et al.
Veröffentlicht: (2026)
von: Khayatan, Pegah, et al.
Veröffentlicht: (2026)
The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models
von: Serra, Alessandro Pietro, et al.
Veröffentlicht: (2024)
von: Serra, Alessandro Pietro, et al.
Veröffentlicht: (2024)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
Multimodal Language Models See Better When They Look Shallower
von: Chen, Haoran, et al.
Veröffentlicht: (2025)
von: Chen, Haoran, et al.
Veröffentlicht: (2025)
Head Pursuit: Probing Attention Specialization in Multimodal Transformers
von: Basile, Lorenzo, et al.
Veröffentlicht: (2025)
von: Basile, Lorenzo, et al.
Veröffentlicht: (2025)
Seeing the Abstract: Translating the Abstract Language for Vision Language Models
von: Talon, Davide, et al.
Veröffentlicht: (2025)
von: Talon, Davide, et al.
Veröffentlicht: (2025)
Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
von: Yang, Yi, et al.
Veröffentlicht: (2025)
von: Yang, Yi, et al.
Veröffentlicht: (2025)
Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals
von: Ortu, Francesco, et al.
Veröffentlicht: (2024)
von: Ortu, Francesco, et al.
Veröffentlicht: (2024)
When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models
von: Yan, Yuping, et al.
Veröffentlicht: (2025)
von: Yan, Yuping, et al.
Veröffentlicht: (2025)
Talk in Pieces, See in Whole: Disentangling and Hierarchical Aggregating Representations for Language-based Object Detection
von: An, Sojung, et al.
Veröffentlicht: (2025)
von: An, Sojung, et al.
Veröffentlicht: (2025)
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
von: Saini, Harshvardhan, et al.
Veröffentlicht: (2026)
von: Saini, Harshvardhan, et al.
Veröffentlicht: (2026)
World Models That Know When They Don't Know - Controllable Video Generation with Calibrated Uncertainty
von: Mei, Zhiting, et al.
Veröffentlicht: (2025)
von: Mei, Zhiting, et al.
Veröffentlicht: (2025)
InsightSee: Advancing Multi-agent Vision-Language Models for Enhanced Visual Understanding
von: Zhang, Huaxiang, et al.
Veröffentlicht: (2024)
von: Zhang, Huaxiang, et al.
Veröffentlicht: (2024)
LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
von: Shen, Yuxiang, et al.
Veröffentlicht: (2026)
von: Shen, Yuxiang, et al.
Veröffentlicht: (2026)
SegSub: Evaluating Robustness to Knowledge Conflicts and Hallucinations in Vision-Language Models
von: Carragher, Peter, et al.
Veröffentlicht: (2025)
von: Carragher, Peter, et al.
Veröffentlicht: (2025)
Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models
von: Panchal, Utsav, et al.
Veröffentlicht: (2025)
von: Panchal, Utsav, et al.
Veröffentlicht: (2025)
Do Vision-Language Models See Urban Scenes as People Do? An Urban Perception Benchmark
von: Mushkani, Rashid
Veröffentlicht: (2025)
von: Mushkani, Rashid
Veröffentlicht: (2025)
Bad Seeing or Bad Thinking? Rewarding Perception for Vision-Language Reasoning
von: Wang, Haozhe, et al.
Veröffentlicht: (2026)
von: Wang, Haozhe, et al.
Veröffentlicht: (2026)
When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models
von: Yakun, Cui, et al.
Veröffentlicht: (2026)
von: Yakun, Cui, et al.
Veröffentlicht: (2026)
Does AI See like Art Historians? Interpreting How Vision Language Models Recognize Artistic Style
von: Limpijankit, Marvin, et al.
Veröffentlicht: (2026)
von: Limpijankit, Marvin, et al.
Veröffentlicht: (2026)
Delineating Knowledge Boundaries for Honest Large Vision-Language Models
von: Song, Junru, et al.
Veröffentlicht: (2026)
von: Song, Junru, et al.
Veröffentlicht: (2026)
Mixed Signals: Decoding VLMs' Reasoning and Underlying Bias in Vision-Language Conflict
von: Pezeshkpour, Pouya, et al.
Veröffentlicht: (2025)
von: Pezeshkpour, Pouya, et al.
Veröffentlicht: (2025)
When Visuals Aren't the Problem: Evaluating Vision-Language Models on Misleading Data Visualizations
von: Lalai, Harsh Nishant, et al.
Veröffentlicht: (2026)
von: Lalai, Harsh Nishant, et al.
Veröffentlicht: (2026)
Seeing Like Radiologists: Context- and Gaze-Guided Vision-Language Pretraining for Chest X-rays
von: Liu, Kang, et al.
Veröffentlicht: (2026)
von: Liu, Kang, et al.
Veröffentlicht: (2026)
DualEdit: Dual Editing for Knowledge Updating in Vision-Language Models
von: Shi, Zhiyi, et al.
Veröffentlicht: (2025)
von: Shi, Zhiyi, et al.
Veröffentlicht: (2025)
Knowing When Not to Predict: Self Supervised Learning and Abstention for Safer DR Screening
von: Chopra, Muskaan, et al.
Veröffentlicht: (2026)
von: Chopra, Muskaan, et al.
Veröffentlicht: (2026)
Beyond Words: Multimodal LLM Knows When to Speak
von: Liao, Zikai, et al.
Veröffentlicht: (2025)
von: Liao, Zikai, et al.
Veröffentlicht: (2025)
When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models
von: Lu, Hui, et al.
Veröffentlicht: (2025)
von: Lu, Hui, et al.
Veröffentlicht: (2025)
When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs
von: Fang, Yu, et al.
Veröffentlicht: (2026)
von: Fang, Yu, et al.
Veröffentlicht: (2026)
Seeing Symbols, Missing Cultures: Probing Vision-Language Models' Reasoning on Fire Imagery and Cultural Meaning
von: Yu, Haorui, et al.
Veröffentlicht: (2025)
von: Yu, Haorui, et al.
Veröffentlicht: (2025)
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models
von: He, Zoe Wanying, et al.
Veröffentlicht: (2025)
von: He, Zoe Wanying, et al.
Veröffentlicht: (2025)
Distilling Cross-Modal Knowledge via Feature Disentanglement
von: Liu, Junhong, et al.
Veröffentlicht: (2025)
von: Liu, Junhong, et al.
Veröffentlicht: (2025)
See then Tell: Enhancing Key Information Extraction with Vision Grounding
von: Liu, Shuhang, et al.
Veröffentlicht: (2024)
von: Liu, Shuhang, et al.
Veröffentlicht: (2024)
FloodVision: Urban Flood Depth Estimation Using Foundation Vision-Language Models and Domain Knowledge Graph
von: Liu, Zhangding, et al.
Veröffentlicht: (2025)
von: Liu, Zhangding, et al.
Veröffentlicht: (2025)
When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models
von: Hou, Jiacheng, et al.
Veröffentlicht: (2026)
von: Hou, Jiacheng, et al.
Veröffentlicht: (2026)
AI-Generated Images: What Humans and Machines See When They Look at the Same Image
von: Poletti, Silvia, et al.
Veröffentlicht: (2026)
von: Poletti, Silvia, et al.
Veröffentlicht: (2026)
Disentanglement and Compositionality of Letter Identity and Letter Position in Variational Auto-Encoder Vision Models
von: Bianchi, Bruno, et al.
Veröffentlicht: (2024)
von: Bianchi, Bruno, et al.
Veröffentlicht: (2024)
Unifying Biomedical Vision-Language Expertise: Towards a Generalist Foundation Model via Multi-CLIP Knowledge Distillation
von: Wang, Shansong, et al.
Veröffentlicht: (2025)
von: Wang, Shansong, et al.
Veröffentlicht: (2025)
MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues
von: Du, Dazhao, et al.
Veröffentlicht: (2026)
von: Du, Dazhao, et al.
Veröffentlicht: (2026)
Revisiting KRISP: A Lightweight Reproduction and Analysis of Knowledge-Enhanced Vision-Language Models
von: Dutta, Souradeep, et al.
Veröffentlicht: (2025)
von: Dutta, Souradeep, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs
von: Khayatan, Pegah, et al.
Veröffentlicht: (2026) -
The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models
von: Serra, Alessandro Pietro, et al.
Veröffentlicht: (2024) -
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
von: Zhang, Yue, et al.
Veröffentlicht: (2026) -
Multimodal Language Models See Better When They Look Shallower
von: Chen, Haoran, et al.
Veröffentlicht: (2025) -
Head Pursuit: Probing Attention Specialization in Multimodal Transformers
von: Basile, Lorenzo, et al.
Veröffentlicht: (2025)