SportSkills: Physical Skill Learning from Sports Instructional Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ashutosh, Kumar, Wu, Chi Hsuan, Grauman, Kristen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917362832244736
author Ashutosh, Kumar
Wu, Chi Hsuan
Grauman, Kristen
author_facet Ashutosh, Kumar
Wu, Chi Hsuan
Grauman, Kristen
contents Current large-scale video datasets focus on general human activity, but lack depth of coverage on fine-grained activities needed to address physical skill learning. We introduce SportSkills, the first large-scale sports dataset geared towards physical skill learning with in-the-wild video. SportSkills has more than 360k instructional videos containing more than 630k visual demonstrations paired with instructional narrations explaining the know-how behind the actions from 55 varied sports. Through a suite of experiments, we show that SportSkills unlocks the ability to understand fine-grained differences between physical actions. Our representation achieves gains of up to 4x with the same model trained on traditional activity-centric datasets. Crucially, building on SportSkills, we introduce the first large-scale task formulation of mistake-conditioned instructional video retrieval, bridging representation learning and actionable feedback generation (e.g., "here's my execution of a skill; which video clip should I watch to improve it?"). Formal evaluations by professional coaches show our retrieval approach significantly advances the ability of video models to personalize visual instructions for a user query.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25163
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SportSkills: Physical Skill Learning from Sports Instructional Videos
Ashutosh, Kumar
Wu, Chi Hsuan
Grauman, Kristen
Computer Vision and Pattern Recognition
Current large-scale video datasets focus on general human activity, but lack depth of coverage on fine-grained activities needed to address physical skill learning. We introduce SportSkills, the first large-scale sports dataset geared towards physical skill learning with in-the-wild video. SportSkills has more than 360k instructional videos containing more than 630k visual demonstrations paired with instructional narrations explaining the know-how behind the actions from 55 varied sports. Through a suite of experiments, we show that SportSkills unlocks the ability to understand fine-grained differences between physical actions. Our representation achieves gains of up to 4x with the same model trained on traditional activity-centric datasets. Crucially, building on SportSkills, we introduce the first large-scale task formulation of mistake-conditioned instructional video retrieval, bridging representation learning and actionable feedback generation (e.g., "here's my execution of a skill; which video clip should I watch to improve it?"). Formal evaluations by professional coaches show our retrieval approach significantly advances the ability of video models to personalize visual instructions for a user query.
title SportSkills: Physical Skill Learning from Sports Instructional Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.25163