ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rawte, Vipula, Jain, Sarthak, Sinha, Aarush, Kaushik, Garv, Bansal, Aman, Vishwanath, Prathiksha Rumale, Jain, Samyak Rajesh, Reganti, Aishwarya Naresh, Jain, Vinija, Chadha, Aman, Sheth, Amit P., Das, Amitava
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908275773014016
author Rawte, Vipula
Jain, Sarthak
Sinha, Aarush
Kaushik, Garv
Bansal, Aman
Vishwanath, Prathiksha Rumale
Jain, Samyak Rajesh
Reganti, Aishwarya Naresh
Jain, Vinija
Chadha, Aman
Sheth, Amit P.
Das, Amitava
author_facet Rawte, Vipula
Jain, Sarthak
Sinha, Aarush
Kaushik, Garv
Bansal, Aman
Vishwanath, Prathiksha Rumale
Jain, Samyak Rajesh
Reganti, Aishwarya Naresh
Jain, Vinija
Chadha, Aman
Sheth, Amit P.
Das, Amitava
contents Recent advances in Large Multimodal Models (LMMs) have expanded their capabilities to video understanding, with Text-to-Video (T2V) models excelling in generating videos from textual prompts. However, they still frequently produce hallucinated content, revealing AI-generated inconsistencies. We introduce ViBe (https://vibe-t2v-bench.github.io/): a large-scale dataset of hallucinated videos from open-source T2V models. We identify five major hallucination types: Vanishing Subject, Omission Error, Numeric Variability, Subject Dysmorphia, and Visual Incongruity. Using ten T2V models, we generated and manually annotated 3,782 videos from 837 diverse MS COCO captions. Our proposed benchmark includes a dataset of hallucinated videos and a classification framework using video embeddings. ViBe serves as a critical resource for evaluating T2V reliability and advancing hallucination detection. We establish classification as a baseline, with the TimeSFormer + CNN ensemble achieving the best performance (0.345 accuracy, 0.342 F1 score). While initial baselines proposed achieve modest accuracy, this highlights the difficulty of automated hallucination detection and the need for improved methods. Our research aims to drive the development of more robust T2V models and evaluate their outputs based on user preferences.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10867
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models
Rawte, Vipula
Jain, Sarthak
Sinha, Aarush
Kaushik, Garv
Bansal, Aman
Vishwanath, Prathiksha Rumale
Jain, Samyak Rajesh
Reganti, Aishwarya Naresh
Jain, Vinija
Chadha, Aman
Sheth, Amit P.
Das, Amitava
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent advances in Large Multimodal Models (LMMs) have expanded their capabilities to video understanding, with Text-to-Video (T2V) models excelling in generating videos from textual prompts. However, they still frequently produce hallucinated content, revealing AI-generated inconsistencies. We introduce ViBe (https://vibe-t2v-bench.github.io/): a large-scale dataset of hallucinated videos from open-source T2V models. We identify five major hallucination types: Vanishing Subject, Omission Error, Numeric Variability, Subject Dysmorphia, and Visual Incongruity. Using ten T2V models, we generated and manually annotated 3,782 videos from 837 diverse MS COCO captions. Our proposed benchmark includes a dataset of hallucinated videos and a classification framework using video embeddings. ViBe serves as a critical resource for evaluating T2V reliability and advancing hallucination detection. We establish classification as a baseline, with the TimeSFormer + CNN ensemble achieving the best performance (0.345 accuracy, 0.342 F1 score). While initial baselines proposed achieve modest accuracy, this highlights the difficulty of automated hallucination detection and the need for improved methods. Our research aims to drive the development of more robust T2V models and evaluate their outputs based on user preferences.
title ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.10867