Domain Adaptation of VLM for Soccer Video Understanding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jiang, Tiancheng, Wang, Henry, Salekin, Md Sirajus, Atighehchian, Parmida, Zhang, Shinan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916828536635392
author Jiang, Tiancheng
Wang, Henry
Salekin, Md Sirajus
Atighehchian, Parmida
Zhang, Shinan
author_facet Jiang, Tiancheng
Wang, Henry
Salekin, Md Sirajus
Atighehchian, Parmida
Zhang, Shinan
contents Vision Language Models (VLMs) have demonstrated strong performance in multi-modal tasks by effectively aligning visual and textual representations. However, most video understanding VLM research has been domain-agnostic, leaving the understanding of their transfer learning capability to specialized domains under-explored. In this work, we address this by exploring the adaptability of open-source VLMs to specific domains, and focusing on soccer as an initial case study. Our approach uses large-scale soccer datasets and LLM to create instruction-following data, and use them to iteratively fine-tune the general-domain VLM in a curriculum learning fashion (first teaching the model key soccer concepts to then question answering tasks). The final adapted model, trained using a curated dataset of 20k video clips, exhibits significant improvement in soccer-specific tasks compared to the base model, with a 37.5% relative improvement for the visual question-answering task and an accuracy improvement from 11.8% to 63.5% for the downstream soccer action classification task.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13860
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Domain Adaptation of VLM for Soccer Video Understanding
Jiang, Tiancheng
Wang, Henry
Salekin, Md Sirajus
Atighehchian, Parmida
Zhang, Shinan
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision Language Models (VLMs) have demonstrated strong performance in multi-modal tasks by effectively aligning visual and textual representations. However, most video understanding VLM research has been domain-agnostic, leaving the understanding of their transfer learning capability to specialized domains under-explored. In this work, we address this by exploring the adaptability of open-source VLMs to specific domains, and focusing on soccer as an initial case study. Our approach uses large-scale soccer datasets and LLM to create instruction-following data, and use them to iteratively fine-tune the general-domain VLM in a curriculum learning fashion (first teaching the model key soccer concepts to then question answering tasks). The final adapted model, trained using a curated dataset of 20k video clips, exhibits significant improvement in soccer-specific tasks compared to the base model, with a 37.5% relative improvement for the visual question-answering task and an accuracy improvement from 11.8% to 63.5% for the downstream soccer action classification task.
title Domain Adaptation of VLM for Soccer Video Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.13860