Saved in:
| Main Authors: | Ling, Chen, Zhang, Tongwei, Li, Hanqian, Ding, Nai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.07737 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
by: Liu, Zhining, et al.
Published: (2025)
by: Liu, Zhining, et al.
Published: (2025)
Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models
by: Li, Nanxi, et al.
Published: (2026)
by: Li, Nanxi, et al.
Published: (2026)
Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models
by: Panchal, Utsav, et al.
Published: (2025)
by: Panchal, Utsav, et al.
Published: (2025)
Believing is Seeing: Unobserved Object Detection using Generative Models
by: Bhattacharjee, Subhransu S., et al.
Published: (2024)
by: Bhattacharjee, Subhransu S., et al.
Published: (2024)
Seeing is Believing: Robust Vision-Guided Cross-Modal Prompt Learning under Label Noise
by: Geng, Zibin, et al.
Published: (2026)
by: Geng, Zibin, et al.
Published: (2026)
Seeing Beyond the Scene: Analyzing and Mitigating Background Bias in Action Recognition
by: Zhou, Ellie, et al.
Published: (2025)
by: Zhou, Ellie, et al.
Published: (2025)
Seeing Isn't Believing: Context-Aware Adversarial Patch Synthesis via Conditional GAN
by: Kazoom, Roie, et al.
Published: (2025)
by: Kazoom, Roie, et al.
Published: (2025)
Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
by: Ou, Siqu, et al.
Published: (2026)
by: Ou, Siqu, et al.
Published: (2026)
CounterCount: A Diagnostic Framework for Counting Bias in Vision Language Models
by: Alzahrani, Reem, et al.
Published: (2026)
by: Alzahrani, Reem, et al.
Published: (2026)
Seeing Right but Saying Wrong: Inter- and Intra-Layer Refinement in MLLMs without Training
by: Song, Shezheng, et al.
Published: (2026)
by: Song, Shezheng, et al.
Published: (2026)
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
by: Deng, Ailin, et al.
Published: (2024)
by: Deng, Ailin, et al.
Published: (2024)
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
by: Qiu, Yansheng, et al.
Published: (2025)
by: Qiu, Yansheng, et al.
Published: (2025)
Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs
by: Huang, Jincai, et al.
Published: (2026)
by: Huang, Jincai, et al.
Published: (2026)
Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization
by: Chen, Yuqi, et al.
Published: (2026)
by: Chen, Yuqi, et al.
Published: (2026)
RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding
by: Yang, Jihan, et al.
Published: (2023)
by: Yang, Jihan, et al.
Published: (2023)
Do Vision-Language Models See Urban Scenes as People Do? An Urban Perception Benchmark
by: Mushkani, Rashid
Published: (2025)
by: Mushkani, Rashid
Published: (2025)
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
by: Guo, Pinxue, et al.
Published: (2025)
by: Guo, Pinxue, et al.
Published: (2025)
SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Model
by: Shi, Yukai, et al.
Published: (2025)
by: Shi, Yukai, et al.
Published: (2025)
UltraLED: Learning to See Everything in Ultra-High Dynamic Range Scenes
by: Meng, Yuang, et al.
Published: (2025)
by: Meng, Yuang, et al.
Published: (2025)
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks
by: Zhou, Ting, et al.
Published: (2024)
by: Zhou, Ting, et al.
Published: (2024)
Intuitive Axial Augmentation Using Polar-Sine-Based Piecewise Distortion for Medical Slice-Wise Segmentation
by: Zhang, Yiqin, et al.
Published: (2024)
by: Zhang, Yiqin, et al.
Published: (2024)
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
Dense Connector for MLLMs
by: Yao, Huanjin, et al.
Published: (2024)
by: Yao, Huanjin, et al.
Published: (2024)
See the past: Time-Reversed Scene Reconstruction from Thermal Traces Using Visual Language Models
by: Contreras, Kebin, et al.
Published: (2025)
by: Contreras, Kebin, et al.
Published: (2025)
ST-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with MLLMs
by: Luo, Bingjun, et al.
Published: (2026)
by: Luo, Bingjun, et al.
Published: (2026)
Evaluating Modern Approaches in 3D Scene Reconstruction: NeRF vs Gaussian-Based Methods
by: Zhou, Yiming, et al.
Published: (2024)
by: Zhou, Yiming, et al.
Published: (2024)
OpenBias: Open-set Bias Detection in Text-to-Image Generative Models
by: D'Incà, Moreno, et al.
Published: (2024)
by: D'Incà, Moreno, et al.
Published: (2024)
The Scene Language: Representing Scenes with Programs, Words, and Embeddings
by: Zhang, Yunzhi, et al.
Published: (2024)
by: Zhang, Yunzhi, et al.
Published: (2024)
A Visual Semantic Adaptive Watermark grounded by Prefix-Tuning for Large Vision-Language Model
by: Zheng, Qi, et al.
Published: (2026)
by: Zheng, Qi, et al.
Published: (2026)
CrystaL: Spontaneous Emergence of Visual Latents in MLLMs
by: Zhang, Yang, et al.
Published: (2026)
by: Zhang, Yang, et al.
Published: (2026)
LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference
by: Yuan, Jianhao, et al.
Published: (2025)
by: Yuan, Jianhao, et al.
Published: (2025)
Open-Sora Plan: Open-Source Large Video Generation Model
by: Lin, Bin, et al.
Published: (2024)
by: Lin, Bin, et al.
Published: (2024)
Weakly Supervised Point Clouds Transformer for 3D Object Detection
by: Tang, Zuojin, et al.
Published: (2023)
by: Tang, Zuojin, et al.
Published: (2023)
SceneLLM: Implicit Language Reasoning in LLM for Dynamic Scene Graph Generation
by: Zhang, Hang, et al.
Published: (2024)
by: Zhang, Hang, et al.
Published: (2024)
TOPO-Bench: An Open-Source Topological Mapping Evaluation Framework with Quantifiable Perceptual Aliasing
by: Wang, Jiaming, et al.
Published: (2025)
by: Wang, Jiaming, et al.
Published: (2025)
VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model
by: Wang, Sibo, et al.
Published: (2024)
by: Wang, Sibo, et al.
Published: (2024)
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators
by: Zhou, Jiazhou, et al.
Published: (2026)
by: Zhou, Jiazhou, et al.
Published: (2026)
Multimodal Continual Learning with MLLMs from Multi-scenario Perspectives
by: Jiang, Kai, et al.
Published: (2025)
by: Jiang, Kai, et al.
Published: (2025)
See in Depth: Training-Free Surgical Scene Segmentation with Monocular Depth Priors
by: Yang, Kunyi, et al.
Published: (2025)
by: Yang, Kunyi, et al.
Published: (2025)
ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image Intermediary
by: Gu, Zeqi, et al.
Published: (2025)
by: Gu, Zeqi, et al.
Published: (2025)
Similar Items
-
Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
by: Liu, Zhining, et al.
Published: (2025) -
Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models
by: Li, Nanxi, et al.
Published: (2026) -
Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models
by: Panchal, Utsav, et al.
Published: (2025) -
Believing is Seeing: Unobserved Object Detection using Generative Models
by: Bhattacharjee, Subhransu S., et al.
Published: (2024) -
Seeing is Believing: Robust Vision-Guided Cross-Modal Prompt Learning under Label Noise
by: Geng, Zibin, et al.
Published: (2026)