TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: You, Ling, Huang, Wenxuan, Xie, Xinni, Wei, Xiangyi, Li, Bangyan, Lin, Shaohui, Li, Yang, Wang, Changbo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918003522666496
author You, Ling
Huang, Wenxuan
Xie, Xinni
Wei, Xiangyi
Li, Bangyan
Lin, Shaohui
Li, Yang
Wang, Changbo
author_facet You, Ling
Huang, Wenxuan
Xie, Xinni
Wei, Xiangyi
Li, Bangyan
Lin, Shaohui
Li, Yang
Wang, Changbo
contents Soccer is a globally popular sporting event, typically characterized by long matches and distinctive highlight moments. Recent advances in Multimodal Large Language Models (MLLMs) offer promising capabilities in temporal grounding and video understanding, soccer commentary generation often requires precise temporal localization and semantically rich descriptions over long-form video. However, existing soccer MLLMs often rely on the temporal a priori for caption generation, so they cannot process the soccer video end-to-end. While some traditional approaches follow a two-step paradigm that is complex and fails to capture the global context to achieve suboptimal performance. To solve the above issues, we present TimeSoccer, the first end-to-end soccer MLLM for Single-anchor Dense Video Captioning (SDVC) in full-match soccer videos. TimeSoccer jointly predicts timestamps and generates captions in a single pass, enabling global context modeling across 45-minute matches. To support long video understanding of soccer matches, we introduce MoFA-Select, a training-free, motion-aware frame compression module that adaptively selects representative frames via a coarse-to-fine strategy, and incorporates complementary training paradigms to strengthen the model's ability to handle long temporal sequences. Extensive experiments demonstrate that our TimeSoccer achieves State-of-The-Art (SoTA) performance on the SDVC task in an end-to-end form, generating high-quality commentary with accurate temporal alignment and strong semantic relevance.
format Preprint
id arxiv_https___arxiv_org_abs_2504_17365
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation
You, Ling
Huang, Wenxuan
Xie, Xinni
Wei, Xiangyi
Li, Bangyan
Lin, Shaohui
Li, Yang
Wang, Changbo
Computer Vision and Pattern Recognition
Computation and Language
Soccer is a globally popular sporting event, typically characterized by long matches and distinctive highlight moments. Recent advances in Multimodal Large Language Models (MLLMs) offer promising capabilities in temporal grounding and video understanding, soccer commentary generation often requires precise temporal localization and semantically rich descriptions over long-form video. However, existing soccer MLLMs often rely on the temporal a priori for caption generation, so they cannot process the soccer video end-to-end. While some traditional approaches follow a two-step paradigm that is complex and fails to capture the global context to achieve suboptimal performance. To solve the above issues, we present TimeSoccer, the first end-to-end soccer MLLM for Single-anchor Dense Video Captioning (SDVC) in full-match soccer videos. TimeSoccer jointly predicts timestamps and generates captions in a single pass, enabling global context modeling across 45-minute matches. To support long video understanding of soccer matches, we introduce MoFA-Select, a training-free, motion-aware frame compression module that adaptively selects representative frames via a coarse-to-fine strategy, and incorporates complementary training paradigms to strengthen the model's ability to handle long temporal sequences. Extensive experiments demonstrate that our TimeSoccer achieves State-of-The-Art (SoTA) performance on the SDVC task in an end-to-end form, generating high-quality commentary with accurate temporal alignment and strong semantic relevance.
title TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2504.17365