SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xia, Haotian, Ge, Haonan, Zou, Junbo, Choi, Hyun Woo, Zhang, Xuebin, Suradja, Danny, Rui, Botao, Tran, Ethan, Jin, Wendy, Ye, Zhen, Lin, Xiyang, Lai, Christopher, Zhang, Shengjie, Miao, Junwen, Chen, Shichao, Tracy, Rhys, Ordonez, Vicente, Shen, Weining, Chen, Hanjie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918365487955968
author Xia, Haotian
Ge, Haonan
Zou, Junbo
Choi, Hyun Woo
Zhang, Xuebin
Suradja, Danny
Rui, Botao
Tran, Ethan
Jin, Wendy
Ye, Zhen
Lin, Xiyang
Lai, Christopher
Zhang, Shengjie
Miao, Junwen
Chen, Shichao
Tracy, Rhys
Ordonez, Vicente
Shen, Weining
Chen, Hanjie
author_facet Xia, Haotian
Ge, Haonan
Zou, Junbo
Choi, Hyun Woo
Zhang, Xuebin
Suradja, Danny
Rui, Botao
Tran, Ethan
Jin, Wendy
Ye, Zhen
Lin, Xiyang
Lai, Christopher
Zhang, Shengjie
Miao, Junwen
Chen, Shichao
Tracy, Rhys
Ordonez, Vicente
Shen, Weining
Chen, Hanjie
contents Deeply understanding sports requires an intricate blend of fine-grained visual perception and rule-based reasoning - a challenge that pushes the limits of current multimodal models. To succeed, models must master three critical capabilities: perceiving nuanced visual details, applying abstract sport rule knowledge, and grounding that knowledge in specific visual evidence. Current sports benchmarks either cover single sports or lack the detailed reasoning chains and precise visual grounding needed to robustly evaluate these core capabilities in a multi-sport context. To address this gap, we introduce SportR, the first multi-sports large-scale benchmark designed to train and evaluate MLLMs on the fundamental reasoning required for sports intelligence. Our benchmark provides a dataset of 4,789 images and 2,052 videos. To enable granular evaluation, we structure our benchmark around a progressive hierarchy of question-answer pairs designed to probe reasoning at increasing depths - from simple infraction identification to complex penalty prediction. For the most advanced tasks requiring multi-step reasoning, such as determining penalties or explaining tactics, we provide 6,841 high-quality, human-authored Chain of Thought annotations. In addition, our benchmark incorporates both image and video modalities and provides manual bounding box annotations to test visual grounding in the image part directly. Extensive experiments demonstrate the profound difficulty of our benchmark. State-of-the-art baseline models perform poorly on our most challenging tasks. While training on our data via Supervised Fine-Tuning and Reinforcement Learning improves these scores, they remain relatively low, highlighting a significant gap in current model capabilities. SportR presents a new challenge for the community, providing a critical resource to drive future research in multimodal sports reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2511_06499
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports
Xia, Haotian
Ge, Haonan
Zou, Junbo
Choi, Hyun Woo
Zhang, Xuebin
Suradja, Danny
Rui, Botao
Tran, Ethan
Jin, Wendy
Ye, Zhen
Lin, Xiyang
Lai, Christopher
Zhang, Shengjie
Miao, Junwen
Chen, Shichao
Tracy, Rhys
Ordonez, Vicente
Shen, Weining
Chen, Hanjie
Computer Vision and Pattern Recognition
Deeply understanding sports requires an intricate blend of fine-grained visual perception and rule-based reasoning - a challenge that pushes the limits of current multimodal models. To succeed, models must master three critical capabilities: perceiving nuanced visual details, applying abstract sport rule knowledge, and grounding that knowledge in specific visual evidence. Current sports benchmarks either cover single sports or lack the detailed reasoning chains and precise visual grounding needed to robustly evaluate these core capabilities in a multi-sport context. To address this gap, we introduce SportR, the first multi-sports large-scale benchmark designed to train and evaluate MLLMs on the fundamental reasoning required for sports intelligence. Our benchmark provides a dataset of 4,789 images and 2,052 videos. To enable granular evaluation, we structure our benchmark around a progressive hierarchy of question-answer pairs designed to probe reasoning at increasing depths - from simple infraction identification to complex penalty prediction. For the most advanced tasks requiring multi-step reasoning, such as determining penalties or explaining tactics, we provide 6,841 high-quality, human-authored Chain of Thought annotations. In addition, our benchmark incorporates both image and video modalities and provides manual bounding box annotations to test visual grounding in the image part directly. Extensive experiments demonstrate the profound difficulty of our benchmark. State-of-the-art baseline models perform poorly on our most challenging tasks. While training on our data via Supervised Fine-Tuning and Reinforcement Learning improves these scores, they remain relatively low, highlighting a significant gap in current model capabilities. SportR presents a new challenge for the community, providing a critical resource to drive future research in multimodal sports reasoning.
title SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.06499