Beyond Logit Lens: Contextual Embeddings for Robust Hallucination Detection & Grounding in VLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Phukan, Anirudh, Divyansh, Morj, Harshit Kumar, Vaishnavi, Saxena, Apoorv, Goswami, Koustava
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916619957043200
author Phukan, Anirudh
Divyansh
Morj, Harshit Kumar
Vaishnavi
Saxena, Apoorv
Goswami, Koustava
author_facet Phukan, Anirudh
Divyansh
Morj, Harshit Kumar
Vaishnavi
Saxena, Apoorv
Goswami, Koustava
contents The rapid development of Large Multimodal Models (LMMs) has significantly advanced multimodal understanding by harnessing the language abilities of Large Language Models (LLMs) and integrating modality-specific encoders. However, LMMs are plagued by hallucinations that limit their reliability and adoption. While traditional methods to detect and mitigate these hallucinations often involve costly training or rely heavily on external models, recent approaches utilizing internal model features present a promising alternative. In this paper, we critically assess the limitations of the state-of-the-art training-free technique, the logit lens, in handling generalized visual hallucinations. We introduce ContextualLens, a refined method that leverages contextual token embeddings from middle layers of LMMs. This approach significantly improves hallucination detection and grounding across diverse categories, including actions and OCR, while also excelling in tasks requiring contextual understanding, such as spatial relations and attribute comparison. Our novel grounding technique yields highly precise bounding boxes, facilitating a transition from Zero-Shot Object Segmentation to Grounded Visual Question Answering. Our contributions pave the way for more reliable and interpretable multimodal models.
format Preprint
id arxiv_https___arxiv_org_abs_2411_19187
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Beyond Logit Lens: Contextual Embeddings for Robust Hallucination Detection & Grounding in VLMs
Phukan, Anirudh
Divyansh
Morj, Harshit Kumar
Vaishnavi
Saxena, Apoorv
Goswami, Koustava
Computation and Language
The rapid development of Large Multimodal Models (LMMs) has significantly advanced multimodal understanding by harnessing the language abilities of Large Language Models (LLMs) and integrating modality-specific encoders. However, LMMs are plagued by hallucinations that limit their reliability and adoption. While traditional methods to detect and mitigate these hallucinations often involve costly training or rely heavily on external models, recent approaches utilizing internal model features present a promising alternative. In this paper, we critically assess the limitations of the state-of-the-art training-free technique, the logit lens, in handling generalized visual hallucinations. We introduce ContextualLens, a refined method that leverages contextual token embeddings from middle layers of LMMs. This approach significantly improves hallucination detection and grounding across diverse categories, including actions and OCR, while also excelling in tasks requiring contextual understanding, such as spatial relations and attribute comparison. Our novel grounding technique yields highly precise bounding boxes, facilitating a transition from Zero-Shot Object Segmentation to Grounded Visual Question Answering. Our contributions pave the way for more reliable and interpretable multimodal models.
title Beyond Logit Lens: Contextual Embeddings for Robust Hallucination Detection & Grounding in VLMs
topic Computation and Language
url https://arxiv.org/abs/2411.19187