RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dong, Linfeng, Yang, Yuchen, Wu, Hao, Wang, Wei, Hou, Yuenan, Zhong, Zhihang, Sun, Xiao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910003531612160
author Dong, Linfeng
Yang, Yuchen
Wu, Hao
Wang, Wei
Hou, Yuenan
Zhong, Zhihang
Sun, Xiao
author_facet Dong, Linfeng
Yang, Yuchen
Wu, Hao
Wang, Wei
Hou, Yuenan
Zhong, Zhihang
Sun, Xiao
contents We introduce RacketVision, a novel dataset and benchmark for advancing computer vision in sports analytics, covering table tennis, tennis, and badminton. The dataset is the first to provide large-scale, fine-grained annotations for racket pose alongside traditional ball positions, enabling research into complex human-object interactions. It is designed to tackle three interconnected tasks: fine-grained ball tracking, articulated racket pose estimation, and predictive ball trajectory forecasting. Our evaluation of established baselines reveals a critical insight for multi-modal fusion: while naively concatenating racket pose features degrades performance, a CrossAttention mechanism is essential to unlock their value, leading to trajectory prediction results that surpass strong unimodal baselines. RacketVision provides a versatile resource and a strong starting point for future research in dynamic object tracking, conditional motion forecasting, and multimodal analysis in sports. Project page at https://github.com/OrcustD/RacketVision
format Preprint
id arxiv_https___arxiv_org_abs_2511_17045
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis
Dong, Linfeng
Yang, Yuchen
Wu, Hao
Wang, Wei
Hou, Yuenan
Zhong, Zhihang
Sun, Xiao
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
We introduce RacketVision, a novel dataset and benchmark for advancing computer vision in sports analytics, covering table tennis, tennis, and badminton. The dataset is the first to provide large-scale, fine-grained annotations for racket pose alongside traditional ball positions, enabling research into complex human-object interactions. It is designed to tackle three interconnected tasks: fine-grained ball tracking, articulated racket pose estimation, and predictive ball trajectory forecasting. Our evaluation of established baselines reveals a critical insight for multi-modal fusion: while naively concatenating racket pose features degrades performance, a CrossAttention mechanism is essential to unlock their value, leading to trajectory prediction results that surpass strong unimodal baselines. RacketVision provides a versatile resource and a strong starting point for future research in dynamic object tracking, conditional motion forecasting, and multimodal analysis in sports. Project page at https://github.com/OrcustD/RacketVision
title RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2511.17045