Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies
Fuente:
arXiv
Saved in:
| Main Authors: | Hou, Wenjin, Liu, Wei, Hu, Han, Sun, Xiaoxiao, Yeung-Levy, Serena, Fan, Hehe |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Incentivizing Generative Zero-Shot Learning via Outcome-Reward Reinforcement Learning with Visual Cues
by: Hou, Wenjin, et al.
Published: (2026)
by: Hou, Wenjin, et al.
Published: (2026)
Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual Illusions
by: Sun, Xiaoxiao, et al.
Published: (2026)
by: Sun, Xiaoxiao, et al.
Published: (2026)
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
by: He, Zhentao, et al.
Published: (2025)
by: He, Zhentao, et al.
Published: (2025)
Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
by: Endo, Mark, et al.
Published: (2025)
by: Endo, Mark, et al.
Published: (2025)
Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration
by: Endo, Mark, et al.
Published: (2024)
by: Endo, Mark, et al.
Published: (2024)
Seeing is Believing? Enhancing Vision-Language Navigation using Visual Perturbations
by: Zhang, Xuesong, et al.
Published: (2024)
by: Zhang, Xuesong, et al.
Published: (2024)
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
by: Aklilu, Josiah, et al.
Published: (2024)
by: Aklilu, Josiah, et al.
Published: (2024)
Prompt-Aware Adapter: Towards Learning Adaptive Visual Tokens for Multimodal Large Language Models
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
ZeroMamba: Exploring Visual State Space Model for Zero-Shot Learning
by: Hou, Wenjin, et al.
Published: (2024)
by: Hou, Wenjin, et al.
Published: (2024)
IllusionBench+: A Large-scale and Comprehensive Benchmark for Visual Illusion Understanding in Vision-Language Models
by: Zhang, Yiming, et al.
Published: (2025)
by: Zhang, Yiming, et al.
Published: (2025)
Continuous Perception Matters: Diagnosing Temporal Integration Failures in Multimodal Models
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
The Impact of Image Resolution on Biomedical Multimodal Large Language Models
by: Chen, Liangyu, et al.
Published: (2025)
by: Chen, Liangyu, et al.
Published: (2025)
Deepfake Detection Generalization with Diffusion Noise
by: Qi, Hongyuan, et al.
Published: (2026)
by: Qi, Hongyuan, et al.
Published: (2026)
SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
by: Deng, Andong, et al.
Published: (2025)
by: Deng, Andong, et al.
Published: (2025)
Seeing the Evidence, Missing the Answer: Tool-Guided Vision-Language Models on Visual Illusions
by: Wang, Xuesong, et al.
Published: (2026)
by: Wang, Xuesong, et al.
Published: (2026)
Foundation Models Secretly Understand Neural Network Weights: Enhancing Hypernetwork Architectures with Foundation Models
by: Gu, Jeffrey, et al.
Published: (2025)
by: Gu, Jeffrey, et al.
Published: (2025)
The Illusion-Illusion: Vision Language Models See Illusions Where There are None
by: Ullman, Tomer
Published: (2024)
by: Ullman, Tomer
Published: (2024)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024)
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024)
Prompt-Aware Controllable Shadow Removal
by: Chen, Kerui, et al.
Published: (2025)
by: Chen, Kerui, et al.
Published: (2025)
Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
by: Liu, Zhining, et al.
Published: (2025)
by: Liu, Zhining, et al.
Published: (2025)
VMAD: Visual-enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection
by: Deng, Huilin, et al.
Published: (2024)
by: Deng, Huilin, et al.
Published: (2024)
Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models
by: Xuan, Weihao, et al.
Published: (2025)
by: Xuan, Weihao, et al.
Published: (2025)
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
by: Sui, Elaine, et al.
Published: (2024)
by: Sui, Elaine, et al.
Published: (2024)
Illusion-Aware Visual Preprocessing and Anti-Illusion Prompting for Classic Illusion Understanding in Vision-Language Models
by: Zha, Junli, et al.
Published: (2026)
by: Zha, Junli, et al.
Published: (2026)
Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2026)
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2026)
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
by: Luo, Gen, et al.
Published: (2025)
by: Luo, Gen, et al.
Published: (2025)
DeforHMR: Vision Transformer with Deformable Cross-Attention for 3D Human Mesh Recovery
by: Heo, Jaewoo, et al.
Published: (2024)
by: Heo, Jaewoo, et al.
Published: (2024)
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
by: Guo, Pinxue, et al.
Published: (2025)
by: Guo, Pinxue, et al.
Published: (2025)
IAD-GPT: Advancing Visual Knowledge in Multimodal Large Language Model for Industrial Anomaly Detection
by: Li, Zewen, et al.
Published: (2025)
by: Li, Zewen, et al.
Published: (2025)
Multi-Human Mesh Recovery with Transformers
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
by: Wang, Xiaohan, et al.
Published: (2024)
by: Wang, Xiaohan, et al.
Published: (2024)
SPR-128K: A New Benchmark for Spatial Plausibility Reasoning with Multimodal Large Language Models
by: Hu, Zhiyuan, et al.
Published: (2025)
by: Hu, Zhiyuan, et al.
Published: (2025)
Seeing is not Believing: An Identity Hider for Human Vision Privacy Protection
by: Wang, Tao, et al.
Published: (2023)
by: Wang, Tao, et al.
Published: (2023)
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
by: Tan, Weiting, et al.
Published: (2025)
by: Tan, Weiting, et al.
Published: (2025)
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
by: Song, Junha, et al.
Published: (2026)
by: Song, Junha, et al.
Published: (2026)
Diffusion-HPC: Synthetic Data Generation for Human Mesh Recovery in Challenging Domains
by: Weng, Zhenzhen, et al.
Published: (2023)
by: Weng, Zhenzhen, et al.
Published: (2023)
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
by: Yang, Jihan, et al.
Published: (2024)
by: Yang, Jihan, et al.
Published: (2024)
Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models
by: Panchal, Utsav, et al.
Published: (2025)
by: Panchal, Utsav, et al.
Published: (2025)
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
by: Shi, Chufan, et al.
Published: (2026)
by: Shi, Chufan, et al.
Published: (2026)
Believing is Seeing: Unobserved Object Detection using Generative Models
by: Bhattacharjee, Subhransu S., et al.
Published: (2024)
by: Bhattacharjee, Subhransu S., et al.
Published: (2024)
Similar Items
-
Incentivizing Generative Zero-Shot Learning via Outcome-Reward Reinforcement Learning with Visual Cues
by: Hou, Wenjin, et al.
Published: (2026) -
Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual Illusions
by: Sun, Xiaoxiao, et al.
Published: (2026) -
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
by: He, Zhentao, et al.
Published: (2025) -
Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
by: Endo, Mark, et al.
Published: (2025) -
Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration
by: Endo, Mark, et al.
Published: (2024)