Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yexin, Liang, Zhengyang, Wang, Yueze, Wu, Xianfeng, Tang, Feilong, He, Muyang, Li, Jian, Liu, Zheng, Yang, Harry, Lim, Sernam, Zhao, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Multimodal Learning from Data-centric Perspective
by: He, Muyang, et al.
Published: (2024)
by: He, Muyang, et al.
Published: (2024)
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
by: Tang, Feilong, et al.
Published: (2025)
by: Tang, Feilong, et al.
Published: (2025)
Superpipeline: A Universal Approach for Reducing GPU Memory Usage in Large Models
by: Abbasi, Reza, et al.
Published: (2024)
by: Abbasi, Reza, et al.
Published: (2024)
Loki: Representation over Architecture for Diffusion-Based Portrait Animation
by: Navard, Pouyan, et al.
Published: (2026)
by: Navard, Pouyan, et al.
Published: (2026)
Model Reveals What to Cache: Profiling-Based Feature Reuse for Video Diffusion Models
by: Ma, Xuran, et al.
Published: (2025)
by: Ma, Xuran, et al.
Published: (2025)
Action Reimagined: Text-to-Pose Video Editing for Dynamic Human Actions
by: Wang, Lan, et al.
Published: (2024)
by: Wang, Lan, et al.
Published: (2024)
SeeClear: Semantic Distillation Enhances Pixel Condensation for Video Super-Resolution
by: Tang, Qi, et al.
Published: (2024)
by: Tang, Qi, et al.
Published: (2024)
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
by: Liu, Yexin, et al.
Published: (2025)
by: Liu, Yexin, et al.
Published: (2025)
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
by: Yuan, Huaying, et al.
Published: (2025)
by: Yuan, Huaying, et al.
Published: (2025)
Learning Latent Proxies for Controllable Single-Image Relighting
by: Zheng, Haoze, et al.
Published: (2026)
by: Zheng, Haoze, et al.
Published: (2026)
See Further When Clear: Curriculum Consistency Model
by: Liu, Yunpeng, et al.
Published: (2024)
by: Liu, Yunpeng, et al.
Published: (2024)
HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling
by: Liu, Xianjie, et al.
Published: (2025)
by: Liu, Xianjie, et al.
Published: (2025)
Temporal Regularization Makes Your Video Generator Stronger
by: Chen, Harold Haodong, et al.
Published: (2025)
by: Chen, Harold Haodong, et al.
Published: (2025)
E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs
by: Liu, Xianjie, et al.
Published: (2026)
by: Liu, Xianjie, et al.
Published: (2026)
Effective Sorting of Fractional Optical Vortex Modes
by: Mao, Zhengyang, et al.
Published: (2024)
by: Mao, Zhengyang, et al.
Published: (2024)
VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
by: Zheng, Mingzhe, et al.
Published: (2025)
by: Zheng, Mingzhe, et al.
Published: (2025)
VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
by: Zheng, Mingzhe, et al.
Published: (2024)
by: Zheng, Mingzhe, et al.
Published: (2024)
Admitting Ignorance Helps the Video Question Answering Models to Answer
by: Li, Haopeng, et al.
Published: (2025)
by: Li, Haopeng, et al.
Published: (2025)
Seeing Clearly without Training: Mitigating Hallucinations in Multimodal LLMs for Remote Sensing
by: Liu, Yi, et al.
Published: (2026)
by: Liu, Yi, et al.
Published: (2026)
A Mystery Cleared Up
by: Weiss, Harry B. (Harry Bischoff)
Published: (1950)
by: Weiss, Harry B. (Harry Bischoff)
Published: (1950)
Retrieval Augmented Question Answering: When Should LLMs Admit Ignorance?
by: Wang, Dingmin, et al.
Published: (2025)
by: Wang, Dingmin, et al.
Published: (2025)
LightGen: Efficient Image Generation through Knowledge Distillation and Direct Preference Optimization
by: Wu, Xianfeng, et al.
Published: (2025)
by: Wu, Xianfeng, et al.
Published: (2025)
SpikeDerain: Unveiling Clear Videos from Rainy Sequences Using Color Spike Streams
by: Liang, Hanwen, et al.
Published: (2025)
by: Liang, Hanwen, et al.
Published: (2025)
Impacts of time‐restricted feeding on middle‐aged and old mice with obesity
by: Yueze Yang, et al.
Published: (2024)
by: Yueze Yang, et al.
Published: (2024)
OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation
by: Liu, Yexin, et al.
Published: (2025)
by: Liu, Yexin, et al.
Published: (2025)
When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
by: Liu, Hongcheng, et al.
Published: (2025)
by: Liu, Hongcheng, et al.
Published: (2025)
Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
by: Lai, Zhengzhao, et al.
Published: (2025)
by: Lai, Zhengzhao, et al.
Published: (2025)
Efficient Multimodal Large Language Models: A Survey
by: Jin, Yizhang, et al.
Published: (2024)
by: Jin, Yizhang, et al.
Published: (2024)
Springs of Ignorance
by: Saadat, Ramin
Published: (2026)
by: Saadat, Ramin
Published: (2026)
TimeLogic Challenge @ CVPR 2026: Strong MLLMs Meet Evidence-Seeking Agents for Temporal-Logic Video Question Answering
by: Xu, Zhaoyang, et al.
Published: (2026)
by: Xu, Zhaoyang, et al.
Published: (2026)
Synthesizing High-Quality Visual Question Answering from Medical Documents with Generator-Verifier LMMs
by: Huang, Xiaoke, et al.
Published: (2025)
by: Huang, Xiaoke, et al.
Published: (2025)
Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generation
by: Zhang, Jiankun, et al.
Published: (2025)
by: Zhang, Jiankun, et al.
Published: (2025)
Large-scale Dataset Pruning with Dynamic Uncertainty
by: He, Muyang, et al.
Published: (2023)
by: He, Muyang, et al.
Published: (2023)
Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
by: Ou, Siqu, et al.
Published: (2026)
by: Ou, Siqu, et al.
Published: (2026)
STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
by: Li, Yun, et al.
Published: (2025)
by: Li, Yun, et al.
Published: (2025)
Geometric Asymmetry in MoE Specialization: Functional Decorrelation and Representational Overlap
by: Liu, Feilong
Published: (2026)
by: Liu, Feilong
Published: (2026)
Rotary Positional Embeddings as Phase Modulation: Theoretical Bounds on the RoPE Base for Long-Context Transformers
by: Liu, Feilong
Published: (2026)
by: Liu, Feilong
Published: (2026)
Mixture-of-Experts as Soft Clustering: A Dual Jacobian-PCA Spectral Geometry Perspective
by: Liu, Feilong
Published: (2026)
by: Liu, Feilong
Published: (2026)
The First Complete Mitochondrial Genome of Common Hedge Blue Acytolepis puspa (Lepidoptera: Lycaenidae), and Comparative Genomic Analysis Within Polyommatinae
by: Muyang Li, et al.
Published: (2026)
by: Muyang Li, et al.
Published: (2026)
LLMs May Perform MCQA by Selecting the Least Incorrect Option
by: Wang, Haochun, et al.
Published: (2024)
by: Wang, Haochun, et al.
Published: (2024)
Similar Items
-
Efficient Multimodal Learning from Data-centric Perspective
by: He, Muyang, et al.
Published: (2024) -
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
by: Tang, Feilong, et al.
Published: (2025) -
Superpipeline: A Universal Approach for Reducing GPU Memory Usage in Large Models
by: Abbasi, Reza, et al.
Published: (2024) -
Loki: Representation over Architecture for Diffusion-Based Portrait Animation
by: Navard, Pouyan, et al.
Published: (2026) -
Model Reveals What to Cache: Profiling-Based Feature Reuse for Video Diffusion Models
by: Ma, Xuran, et al.
Published: (2025)