Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
Fuente:
arXiv
Salvato in:
| Autori principali: | Shi, Chufan, Yang, Cheng, Wu, Yaokang, Jin, Linghao, Shui, Bo, Berg-Kirkpatrick, Taylor, Ma, Xuezhe |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
di: Yang, Cheng, et al.
Pubblicazione: (2026)
di: Yang, Cheng, et al.
Pubblicazione: (2026)
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
di: Gao, Xin, et al.
Pubblicazione: (2026)
di: Gao, Xin, et al.
Pubblicazione: (2026)
Towards Chapter-to-Chapter Context-Aware Literary Translation via Large Language Models
di: Jin, Linghao, et al.
Pubblicazione: (2024)
di: Jin, Linghao, et al.
Pubblicazione: (2024)
Optical Context Compression Is Just (Bad) Autoencoding
di: Lee, Ivan Yee, et al.
Pubblicazione: (2025)
di: Lee, Ivan Yee, et al.
Pubblicazione: (2025)
LLM2: Let Large Language Models Harness System 2 Reasoning
di: Yang, Cheng, et al.
Pubblicazione: (2024)
di: Yang, Cheng, et al.
Pubblicazione: (2024)
The Illusion-Illusion: Vision Language Models See Illusions Where There are None
di: Ullman, Tomer
Pubblicazione: (2024)
di: Ullman, Tomer
Pubblicazione: (2024)
LogogramNLP: Comparing Visual and Textual Representations of Ancient Logographic Writing Systems for NLP
di: Chen, Danlu, et al.
Pubblicazione: (2024)
di: Chen, Danlu, et al.
Pubblicazione: (2024)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
di: Li, Sifan, et al.
Pubblicazione: (2025)
di: Li, Sifan, et al.
Pubblicazione: (2025)
DecoPrompt : Decoding Prompts Reduces Hallucinations when Large Language Models Meet False Premises
di: Xu, Nan, et al.
Pubblicazione: (2024)
di: Xu, Nan, et al.
Pubblicazione: (2024)
Readability $\ne$ Learnability: Rethinking the Role of Simplicity in Training Small Language Models
di: Lee, Ivan, et al.
Pubblicazione: (2025)
di: Lee, Ivan, et al.
Pubblicazione: (2025)
LLM The Genius Paradox: A Linguistic and Math Expert's Struggle with Simple Word-based Counting Problems
di: Xu, Nan, et al.
Pubblicazione: (2024)
di: Xu, Nan, et al.
Pubblicazione: (2024)
MORL-Prompt: An Empirical Analysis of Multi-Objective Reinforcement Learning for Discrete Prompt Optimization
di: Jafari, Yasaman, et al.
Pubblicazione: (2024)
di: Jafari, Yasaman, et al.
Pubblicazione: (2024)
The Format Tax
di: Lee, Ivan Yee, et al.
Pubblicazione: (2026)
di: Lee, Ivan Yee, et al.
Pubblicazione: (2026)
Measuring How (Not Just Whether) VLMs Build Common Ground
di: Imai, Saki, et al.
Pubblicazione: (2025)
di: Imai, Saki, et al.
Pubblicazione: (2025)
To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs
di: Hong, Rui, et al.
Pubblicazione: (2026)
di: Hong, Rui, et al.
Pubblicazione: (2026)
Alt-Text with Context: Improving Accessibility for Images on Twitter
di: Srivatsan, Nikita, et al.
Pubblicazione: (2023)
di: Srivatsan, Nikita, et al.
Pubblicazione: (2023)
Studying the Soupability of Documents in State Space Models
di: Jafari, Yasaman, et al.
Pubblicazione: (2025)
di: Jafari, Yasaman, et al.
Pubblicazione: (2025)
LiFi: Lightweight Controlled Text Generation with Fine-Grained Control Codes
di: Shi, Chufan, et al.
Pubblicazione: (2024)
di: Shi, Chufan, et al.
Pubblicazione: (2024)
Modeling Community Attitude through Reaction Tone: A Human-AI Collaborative Framework for Evaluating LLM Alignment with Linguistic Behaviors in Online Communities
di: Wen, Nuan, et al.
Pubblicazione: (2026)
di: Wen, Nuan, et al.
Pubblicazione: (2026)
[De|Re]constructing VLMs' Reasoning in Counting
di: Alghisi, Simone, et al.
Pubblicazione: (2025)
di: Alghisi, Simone, et al.
Pubblicazione: (2025)
Light-weight Fine-tuning Method for Defending Adversarial Noise in Pre-trained Medical Vision-Language Models
di: Han, Xu, et al.
Pubblicazione: (2024)
di: Han, Xu, et al.
Pubblicazione: (2024)
PatentEdits: Framing Patent Novelty as Textual Entailment
di: Lee, Ryan, et al.
Pubblicazione: (2024)
di: Lee, Ryan, et al.
Pubblicazione: (2024)
Asymmetric Idiosyncrasies in Multimodal Models
di: Tao, Muzi, et al.
Pubblicazione: (2026)
di: Tao, Muzi, et al.
Pubblicazione: (2026)
Translation or Recitation? Calibrating Evaluation Scores for Machine Translation of Extremely Low-Resource Languages
di: Chen, Danlu, et al.
Pubblicazione: (2026)
di: Chen, Danlu, et al.
Pubblicazione: (2026)
Smaller Language Models are Better Black-box Machine-Generated Text Detectors
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2023)
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2023)
What They Saw, Not Just Where They Looked: Semantic Scanpath Similarity via VLMs and NLP metric
di: Kerkouri, Mohamed Amine, et al.
Pubblicazione: (2026)
di: Kerkouri, Mohamed Amine, et al.
Pubblicazione: (2026)
ContextVis: Envision Contextual Learning and Interaction with Generative Models
di: Shui, Bo, et al.
Pubblicazione: (2024)
di: Shui, Bo, et al.
Pubblicazione: (2024)
Constrained Adaptive Rejection Sampling
di: Parys, Paweł, et al.
Pubblicazione: (2025)
di: Parys, Paweł, et al.
Pubblicazione: (2025)
The Tool Illusion: Rethinking Tool Use in Web Agents
di: Lou, Renze, et al.
Pubblicazione: (2026)
di: Lou, Renze, et al.
Pubblicazione: (2026)
Measuring What VLMs Don't Say: Validation Metrics Hide Clinical Terminology Erasure in Radiology Report Generation
di: Parikh, Aditya, et al.
Pubblicazione: (2026)
di: Parikh, Aditya, et al.
Pubblicazione: (2026)
Re-examining Sexism and Misogyny Classification with Annotator Attitudes
di: Jiang, Aiqi, et al.
Pubblicazione: (2024)
di: Jiang, Aiqi, et al.
Pubblicazione: (2024)
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
di: Wu, Jiaying, et al.
Pubblicazione: (2025)
di: Wu, Jiaying, et al.
Pubblicazione: (2025)
From sunblock to softblock: Analyzing the correlates of neology in published writing and on social media
di: Ryskina, Maria, et al.
Pubblicazione: (2026)
di: Ryskina, Maria, et al.
Pubblicazione: (2026)
Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models
di: Khan, Mohammed Safi Ur Rahman, et al.
Pubblicazione: (2026)
di: Khan, Mohammed Safi Ur Rahman, et al.
Pubblicazione: (2026)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
di: Rostamkhani, Mohammadmostafa, et al.
Pubblicazione: (2024)
di: Rostamkhani, Mohammadmostafa, et al.
Pubblicazione: (2024)
Gecko: An Efficient Neural Architecture Inherently Processing Sequences with Arbitrary Lengths
di: Ma, Xuezhe, et al.
Pubblicazione: (2026)
di: Ma, Xuezhe, et al.
Pubblicazione: (2026)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
di: Zhang, Yue, et al.
Pubblicazione: (2026)
di: Zhang, Yue, et al.
Pubblicazione: (2026)
The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
di: Janiak, Denis, et al.
Pubblicazione: (2025)
di: Janiak, Denis, et al.
Pubblicazione: (2025)
See What LLMs Cannot Answer: A Self-Challenge Framework for Uncovering LLM Weaknesses
di: Chen, Yulong, et al.
Pubblicazione: (2024)
di: Chen, Yulong, et al.
Pubblicazione: (2024)
What Prompts Don't Say: Understanding and Managing Underspecification in LLM Prompts
di: Yang, Chenyang, et al.
Pubblicazione: (2025)
di: Yang, Chenyang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
di: Yang, Cheng, et al.
Pubblicazione: (2026) -
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
di: Gao, Xin, et al.
Pubblicazione: (2026) -
Towards Chapter-to-Chapter Context-Aware Literary Translation via Large Language Models
di: Jin, Linghao, et al.
Pubblicazione: (2024) -
Optical Context Compression Is Just (Bad) Autoencoding
di: Lee, Ivan Yee, et al.
Pubblicazione: (2025) -
LLM2: Let Large Language Models Harness System 2 Reasoning
di: Yang, Cheng, et al.
Pubblicazione: (2024)