PolySmart @ TRECVid 2024 Video Captioning (VTT)
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915120728244224 |
|---|---|
| author | Wu, Jiaxin Zhang, Wengyu Wei, Xiao-Yong Li, Qing |
| author_facet | Wu, Jiaxin Zhang, Wengyu Wei, Xiao-Yong Li, Qing |
| contents | In this paper, we present our methods and results for the Video-To-Text (VTT) task at TRECVid 2024, exploring the capabilities of Vision-Language Models (VLMs) like LLaVA and LLaVA-NeXT-Video in generating natural language descriptions for video content. We investigate the impact of fine-tuning VLMs on VTT datasets to enhance description accuracy, contextual relevance, and linguistic consistency. Our analysis reveals that fine-tuning substantially improves the model's ability to produce more detailed and domain-aligned text, bridging the gap between generic VLM tasks and the specialized needs of VTT. Experimental results demonstrate that our fine-tuned model outperforms baseline VLMs across various evaluation metrics, underscoring the importance of domain-specific tuning for complex VTT tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_15509 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | PolySmart @ TRECVid 2024 Video Captioning (VTT) Wu, Jiaxin Zhang, Wengyu Wei, Xiao-Yong Li, Qing Computer Vision and Pattern Recognition Multimedia In this paper, we present our methods and results for the Video-To-Text (VTT) task at TRECVid 2024, exploring the capabilities of Vision-Language Models (VLMs) like LLaVA and LLaVA-NeXT-Video in generating natural language descriptions for video content. We investigate the impact of fine-tuning VLMs on VTT datasets to enhance description accuracy, contextual relevance, and linguistic consistency. Our analysis reveals that fine-tuning substantially improves the model's ability to produce more detailed and domain-aligned text, bridging the gap between generic VLM tasks and the specialized needs of VTT. Experimental results demonstrate that our fine-tuned model outperforms baseline VLMs across various evaluation metrics, underscoring the importance of domain-specific tuning for complex VTT tasks. |
| title | PolySmart @ TRECVid 2024 Video Captioning (VTT) |
| topic | Computer Vision and Pattern Recognition Multimedia |
| url | https://arxiv.org/abs/2412.15509 |