Towards Automatic Soccer Commentary Generation with Knowledge-Enhanced Visual Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Zeyu, Qin, Xiaoyu, Zhou, Songtao, Yun, Kaifeng, Jia, Jia
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908927641255936
author Jin, Zeyu
Qin, Xiaoyu
Zhou, Songtao
Yun, Kaifeng
Jia, Jia
author_facet Jin, Zeyu
Qin, Xiaoyu
Zhou, Songtao
Yun, Kaifeng
Jia, Jia
contents Soccer commentary plays a crucial role in enhancing the soccer game viewing experience for audiences. Previous studies in automatic soccer commentary generation typically adopt an end-to-end method to generate anonymous live text commentary. Such generated commentary is insufficient in the context of real-world live televised commentary, as it contains anonymous entities, context-dependent errors and lacks statistical insights of the game events. To bridge the gap, we propose GameSight, a two-stage model to address soccer commentary generation as a knowledge-enhanced visual reasoning task, enabling live-televised-like knowledgeable commentary with accurate reference to entities (players and teams). GameSight starts by performing visual reasoning to align anonymous entities with fine-grained visual and contextual analysis. Subsequently, the entity-aligned commentary is refined with knowledge by incorporating external historical statistics and iteratively updated internal game state information. Consequently, GameSight improves the player alignment accuracy by 18.5% on SN-Caption-test-align dataset compared to Gemini 2.5-pro. Combined with further knowledge enhancement, GameSight outperforms in segment-level accuracy and commentary quality, as well as game-level contextual relevance and structural composition. We believe that our work paves the way for a more informative and engaging human-centric experience with the AI sports application. Demo Page: https://gamesight2025.github.io/gamesight2025
format Preprint
id arxiv_https___arxiv_org_abs_2604_00057
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Automatic Soccer Commentary Generation with Knowledge-Enhanced Visual Reasoning
Jin, Zeyu
Qin, Xiaoyu
Zhou, Songtao
Yun, Kaifeng
Jia, Jia
Multimedia
Artificial Intelligence
Soccer commentary plays a crucial role in enhancing the soccer game viewing experience for audiences. Previous studies in automatic soccer commentary generation typically adopt an end-to-end method to generate anonymous live text commentary. Such generated commentary is insufficient in the context of real-world live televised commentary, as it contains anonymous entities, context-dependent errors and lacks statistical insights of the game events. To bridge the gap, we propose GameSight, a two-stage model to address soccer commentary generation as a knowledge-enhanced visual reasoning task, enabling live-televised-like knowledgeable commentary with accurate reference to entities (players and teams). GameSight starts by performing visual reasoning to align anonymous entities with fine-grained visual and contextual analysis. Subsequently, the entity-aligned commentary is refined with knowledge by incorporating external historical statistics and iteratively updated internal game state information. Consequently, GameSight improves the player alignment accuracy by 18.5% on SN-Caption-test-align dataset compared to Gemini 2.5-pro. Combined with further knowledge enhancement, GameSight outperforms in segment-level accuracy and commentary quality, as well as game-level contextual relevance and structural composition. We believe that our work paves the way for a more informative and engaging human-centric experience with the AI sports application. Demo Page: https://gamesight2025.github.io/gamesight2025
title Towards Automatic Soccer Commentary Generation with Knowledge-Enhanced Visual Reasoning
topic Multimedia
Artificial Intelligence
url https://arxiv.org/abs/2604.00057