PolySmart @ TRECVid 2024 Video Captioning (VTT)

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wu, Jiaxin, Zhang, Wengyu, Wei, Xiao-Yong, Li, Qing
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915120728244224
author Wu, Jiaxin
Zhang, Wengyu
Wei, Xiao-Yong
Li, Qing
author_facet Wu, Jiaxin
Zhang, Wengyu
Wei, Xiao-Yong
Li, Qing
contents In this paper, we present our methods and results for the Video-To-Text (VTT) task at TRECVid 2024, exploring the capabilities of Vision-Language Models (VLMs) like LLaVA and LLaVA-NeXT-Video in generating natural language descriptions for video content. We investigate the impact of fine-tuning VLMs on VTT datasets to enhance description accuracy, contextual relevance, and linguistic consistency. Our analysis reveals that fine-tuning substantially improves the model's ability to produce more detailed and domain-aligned text, bridging the gap between generic VLM tasks and the specialized needs of VTT. Experimental results demonstrate that our fine-tuned model outperforms baseline VLMs across various evaluation metrics, underscoring the importance of domain-specific tuning for complex VTT tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15509
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PolySmart @ TRECVid 2024 Video Captioning (VTT)
Wu, Jiaxin
Zhang, Wengyu
Wei, Xiao-Yong
Li, Qing
Computer Vision and Pattern Recognition
Multimedia
In this paper, we present our methods and results for the Video-To-Text (VTT) task at TRECVid 2024, exploring the capabilities of Vision-Language Models (VLMs) like LLaVA and LLaVA-NeXT-Video in generating natural language descriptions for video content. We investigate the impact of fine-tuning VLMs on VTT datasets to enhance description accuracy, contextual relevance, and linguistic consistency. Our analysis reveals that fine-tuning substantially improves the model's ability to produce more detailed and domain-aligned text, bridging the gap between generic VLM tasks and the specialized needs of VTT. Experimental results demonstrate that our fine-tuned model outperforms baseline VLMs across various evaluation metrics, underscoring the importance of domain-specific tuning for complex VTT tasks.
title PolySmart @ TRECVid 2024 Video Captioning (VTT)
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2412.15509