A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Sahoo, Pranab, Meharia, Prabhash, Ghosh, Akash, Saha, Sriparna, Jain, Vinija, Chadha, Aman
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917793446756352
author Sahoo, Pranab
Meharia, Prabhash
Ghosh, Akash
Saha, Sriparna
Jain, Vinija
Chadha, Aman
author_facet Sahoo, Pranab
Meharia, Prabhash
Ghosh, Akash
Saha, Sriparna
Jain, Vinija
Chadha, Aman
contents The rapid advancement of foundation models (FMs) across language, image, audio, and video domains has shown remarkable capabilities in diverse tasks. However, the proliferation of FMs brings forth a critical challenge: the potential to generate hallucinated outputs, particularly in high-stakes applications. The tendency of foundation models to produce hallucinated content arguably represents the biggest hindrance to their widespread adoption in real-world scenarios, especially in domains where reliability and accuracy are paramount. This survey paper presents a comprehensive overview of recent developments that aim to identify and mitigate the problem of hallucination in FMs, spanning text, image, video, and audio modalities. By synthesizing recent advancements in detecting and mitigating hallucination across various modalities, the paper aims to provide valuable insights for researchers, developers, and practitioners. Essentially, it establishes a clear framework encompassing definition, taxonomy, and detection strategies for addressing hallucination in multimodal foundation models, laying the foundation for future research in this pivotal area.
format Preprint
id arxiv_https___arxiv_org_abs_2405_09589
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models
Sahoo, Pranab
Meharia, Prabhash
Ghosh, Akash
Saha, Sriparna
Jain, Vinija
Chadha, Aman
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
The rapid advancement of foundation models (FMs) across language, image, audio, and video domains has shown remarkable capabilities in diverse tasks. However, the proliferation of FMs brings forth a critical challenge: the potential to generate hallucinated outputs, particularly in high-stakes applications. The tendency of foundation models to produce hallucinated content arguably represents the biggest hindrance to their widespread adoption in real-world scenarios, especially in domains where reliability and accuracy are paramount. This survey paper presents a comprehensive overview of recent developments that aim to identify and mitigate the problem of hallucination in FMs, spanning text, image, video, and audio modalities. By synthesizing recent advancements in detecting and mitigating hallucination across various modalities, the paper aims to provide valuable insights for researchers, developers, and practitioners. Essentially, it establishes a clear framework encompassing definition, taxonomy, and detection strategies for addressing hallucination in multimodal foundation models, laying the foundation for future research in this pivotal area.
title A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2405.09589