When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Shu, Yan, Lin, Hangui, Liu, Yexin, Zhang, Yan, Zeng, Gangyan, Li, Yan, Zhou, Yu, Lim, Ser-Nam, Yang, Harry, Sebe, Nicu
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909829470093312
author Shu, Yan
Lin, Hangui
Liu, Yexin
Zhang, Yan
Zeng, Gangyan
Li, Yan
Zhou, Yu
Lim, Ser-Nam
Yang, Harry
Sebe, Nicu
author_facet Shu, Yan
Lin, Hangui
Liu, Yexin
Zhang, Yan
Zeng, Gangyan
Li, Yan
Zhou, Yu
Lim, Ser-Nam
Yang, Harry
Sebe, Nicu
contents Large Multimodal Models (LMMs) have achieved impressive progress in visual perception and reasoning. However, when confronted with visually ambiguous or non-semantic scene text, they often struggle to accurately spot and understand the content, frequently generating semantically plausible yet visually incorrect answers, which we refer to as semantic hallucination. In this work, we investigate the underlying causes of semantic hallucination and identify a key finding: Transformer layers in LLM with stronger attention focus on scene text regions are less prone to producing semantic hallucinations. Thus, we propose a training-free semantic hallucination mitigation framework comprising two key components: (1) ZoomText, a coarse-to-fine strategy that identifies potential text regions without external detectors; and (2) Grounded Layer Correction, which adaptively leverages the internal representations from layers less prone to hallucination to guide decoding, correcting hallucinated outputs for non-semantic samples while preserving the semantics of meaningful ones. To enable rigorous evaluation, we introduce TextHalu-Bench, a benchmark of 1,740 samples spanning both semantic and non-semantic cases, with manually curated question answer pairs designed to probe model hallucinations. Extensive experiments demonstrate that our method not only effectively mitigates semantic hallucination but also achieves strong performance on public benchmarks for scene text spotting and understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05551
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
Shu, Yan
Lin, Hangui
Liu, Yexin
Zhang, Yan
Zeng, Gangyan
Li, Yan
Zhou, Yu
Lim, Ser-Nam
Yang, Harry
Sebe, Nicu
Computer Vision and Pattern Recognition
Large Multimodal Models (LMMs) have achieved impressive progress in visual perception and reasoning. However, when confronted with visually ambiguous or non-semantic scene text, they often struggle to accurately spot and understand the content, frequently generating semantically plausible yet visually incorrect answers, which we refer to as semantic hallucination. In this work, we investigate the underlying causes of semantic hallucination and identify a key finding: Transformer layers in LLM with stronger attention focus on scene text regions are less prone to producing semantic hallucinations. Thus, we propose a training-free semantic hallucination mitigation framework comprising two key components: (1) ZoomText, a coarse-to-fine strategy that identifies potential text regions without external detectors; and (2) Grounded Layer Correction, which adaptively leverages the internal representations from layers less prone to hallucination to guide decoding, correcting hallucinated outputs for non-semantic samples while preserving the semantics of meaningful ones. To enable rigorous evaluation, we introduce TextHalu-Bench, a benchmark of 1,740 samples spanning both semantic and non-semantic cases, with manually curated question answer pairs designed to probe model hallucinations. Extensive experiments demonstrate that our method not only effectively mitigates semantic hallucination but also achieves strong performance on public benchmarks for scene text spotting and understanding.
title When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.05551