Do LVLMs Know What They Know? A Systematic Study of Knowledge Boundary Perception in LVLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Zhikai, Ni, Shiyu, Bi, Keping
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915464392736768
author Ding, Zhikai
Ni, Shiyu
Bi, Keping
author_facet Ding, Zhikai
Ni, Shiyu
Bi, Keping
contents Large vision-language models (LVLMs) demonstrate strong visual question answering (VQA) capabilities but are shown to hallucinate. A reliable model should perceive its knowledge boundaries-knowing what it knows and what it does not. This paper investigates LVLMs' perception of their knowledge boundaries by evaluating three types of confidence signals: probabilistic confidence, answer consistency-based confidence, and verbalized confidence. Experiments on three LVLMs across three VQA datasets show that, although LVLMs possess a reasonable perception level, there is substantial room for improvement. Among the three confidences, probabilistic and consistency-based signals are more reliable indicators, while verbalized confidence often leads to overconfidence. To enhance LVLMs' perception, we adapt several established confidence calibration methods from Large Language Models (LLMs) and propose three effective methods. Additionally, we compare LVLMs with their LLM counterparts, finding that jointly processing visual and textual inputs decreases question-answering performance but reduces confidence, resulting in an improved perception level compared to LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2508_19111
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do LVLMs Know What They Know? A Systematic Study of Knowledge Boundary Perception in LVLMs
Ding, Zhikai
Ni, Shiyu
Bi, Keping
Computation and Language
Large vision-language models (LVLMs) demonstrate strong visual question answering (VQA) capabilities but are shown to hallucinate. A reliable model should perceive its knowledge boundaries-knowing what it knows and what it does not. This paper investigates LVLMs' perception of their knowledge boundaries by evaluating three types of confidence signals: probabilistic confidence, answer consistency-based confidence, and verbalized confidence. Experiments on three LVLMs across three VQA datasets show that, although LVLMs possess a reasonable perception level, there is substantial room for improvement. Among the three confidences, probabilistic and consistency-based signals are more reliable indicators, while verbalized confidence often leads to overconfidence. To enhance LVLMs' perception, we adapt several established confidence calibration methods from Large Language Models (LLMs) and propose three effective methods. Additionally, we compare LVLMs with their LLM counterparts, finding that jointly processing visual and textual inputs decreases question-answering performance but reduces confidence, resulting in an improved perception level compared to LLMs.
title Do LVLMs Know What They Know? A Systematic Study of Knowledge Boundary Perception in LVLMs
topic Computation and Language
url https://arxiv.org/abs/2508.19111