FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Jing, Liqiang, Lai, Viet, Yoon, Seunghyun, Bui, Trung, Du, Xinya
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908441692340224
author Jing, Liqiang
Lai, Viet
Yoon, Seunghyun
Bui, Trung
Du, Xinya
author_facet Jing, Liqiang
Lai, Viet
Yoon, Seunghyun
Bui, Trung
Du, Xinya
contents Video Multimodal Large Language Models (VideoMLLMs) have achieved remarkable progress in both Video-to-Text and Text-to-Video tasks. However, they often suffer fro hallucinations, generating content that contradicts the visual input. Existing evaluation methods are limited to one task (e.g., V2T) and also fail to assess hallucinations in open-ended, free-form responses. To address this gap, we propose FIFA, a unified FaIthFulness evAluation framework that extracts comprehensive descriptive facts, models their semantic dependencies via a Spatio-Temporal Semantic Dependency Graph, and verifies them using VideoQA models. We further introduce Post-Correction, a tool-based correction framework that revises hallucinated content. Extensive experiments demonstrate that FIFA aligns more closely with human judgment than existing evaluation methods, and that Post-Correction effectively improves factual consistency in both text and video generation.
format Preprint
id arxiv_https___arxiv_org_abs_2507_06523
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation
Jing, Liqiang
Lai, Viet
Yoon, Seunghyun
Bui, Trung
Du, Xinya
Computer Vision and Pattern Recognition
Computation and Language
Graphics
Video Multimodal Large Language Models (VideoMLLMs) have achieved remarkable progress in both Video-to-Text and Text-to-Video tasks. However, they often suffer fro hallucinations, generating content that contradicts the visual input. Existing evaluation methods are limited to one task (e.g., V2T) and also fail to assess hallucinations in open-ended, free-form responses. To address this gap, we propose FIFA, a unified FaIthFulness evAluation framework that extracts comprehensive descriptive facts, models their semantic dependencies via a Spatio-Temporal Semantic Dependency Graph, and verifies them using VideoQA models. We further introduce Post-Correction, a tool-based correction framework that revises hallucinated content. Extensive experiments demonstrate that FIFA aligns more closely with human judgment than existing evaluation methods, and that Post-Correction effectively improves factual consistency in both text and video generation.
title FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation
topic Computer Vision and Pattern Recognition
Computation and Language
Graphics
url https://arxiv.org/abs/2507.06523