SportSkills: Physical Skill Learning from Sports Instructional Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Ashutosh, Kumar, Wu, Chi Hsuan, Grauman, Kristen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SkillSight: Efficient First-Person Skill Assessment with Gaze
by: Wu, Chi Hsuan, et al.
Published: (2025)
by: Wu, Chi Hsuan, et al.
Published: (2025)
Learning Skill-Attributes for Transferable Assessment in Video
by: Ashutosh, Kumar, et al.
Published: (2025)
by: Ashutosh, Kumar, et al.
Published: (2025)
Stitch-a-Demo: Video Demonstrations from Multistep Descriptions
by: Wu, Chi Hsuan, et al.
Published: (2025)
by: Wu, Chi Hsuan, et al.
Published: (2025)
ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos
by: Somayazulu, Arjun, et al.
Published: (2026)
by: Somayazulu, Arjun, et al.
Published: (2026)
Detours for Navigating Instructional Videos
by: Ashutosh, Kumar, et al.
Published: (2024)
by: Ashutosh, Kumar, et al.
Published: (2024)
Learning Object State Changes in Videos: An Open-World Perspective
by: Xue, Zihui, et al.
Published: (2023)
by: Xue, Zihui, et al.
Published: (2023)
FIction: 4D Future Interaction Prediction from Video
by: Ashutosh, Kumar, et al.
Published: (2024)
by: Ashutosh, Kumar, et al.
Published: (2024)
ExpertAF: Expert Actionable Feedback from Video
by: Ashutosh, Kumar, et al.
Published: (2024)
by: Ashutosh, Kumar, et al.
Published: (2024)
HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
by: An, Joungbin, et al.
Published: (2025)
by: An, Joungbin, et al.
Published: (2025)
SoundingActions: Learning How Actions Sound from Narrated Egocentric Videos
by: Chen, Changan, et al.
Published: (2024)
by: Chen, Changan, et al.
Published: (2024)
SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
PATS: Proficiency-Aware Temporal Sampling for Multi-View Sports Skill Assessment
by: Bianchi, Edoardo, et al.
Published: (2025)
by: Bianchi, Edoardo, et al.
Published: (2025)
UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding
by: An, Joungbin, et al.
Published: (2026)
by: An, Joungbin, et al.
Published: (2026)
Vid2Coach: Transforming How-To Videos into Task Assistants
by: Huh, Mina, et al.
Published: (2025)
by: Huh, Mina, et al.
Published: (2025)
Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video
by: Adebi, Daniel, et al.
Published: (2025)
by: Adebi, Daniel, et al.
Published: (2025)
Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos
by: Majumder, Sagnik, et al.
Published: (2024)
by: Majumder, Sagnik, et al.
Published: (2024)
Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models
by: Baid, Ami, et al.
Published: (2026)
by: Baid, Ami, et al.
Published: (2026)
Progress-Aware Video Frame Captioning
by: Xue, Zihui, et al.
Published: (2024)
by: Xue, Zihui, et al.
Published: (2024)
Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos
by: Majumder, Sagnik, et al.
Published: (2024)
by: Majumder, Sagnik, et al.
Published: (2024)
MistExit: Learning to Exit for Early Mistake Detection in Procedural Videos
by: Majumder, Sagnik, et al.
Published: (2026)
by: Majumder, Sagnik, et al.
Published: (2026)
EgoExo-WM: Unlocking Exo Video for Ego World Models
by: Tran, Danny, et al.
Published: (2026)
by: Tran, Danny, et al.
Published: (2026)
Put Myself in Your Shoes: Lifting the Egocentric Perspective from Exocentric Videos
by: Luo, Mi, et al.
Published: (2024)
by: Luo, Mi, et al.
Published: (2024)
Sports Re-ID: Improving Re-Identification Of Players In Broadcast Videos Of Team Sports
by: Comandur, Bharath
Published: (2022)
by: Comandur, Bharath
Published: (2022)
Sports-QA: A Large-Scale Video Question Answering Benchmark for Complex and Professional Sports
by: Li, Haopeng, et al.
Published: (2024)
by: Li, Haopeng, et al.
Published: (2024)
HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness
by: Xue, Zihui, et al.
Published: (2024)
by: Xue, Zihui, et al.
Published: (2024)
Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos
by: Majumder, Sagnik, et al.
Published: (2023)
by: Majumder, Sagnik, et al.
Published: (2023)
ViSTec: Video Modeling for Sports Technique Recognition and Tactical Analysis
by: He, Yuchen, et al.
Published: (2024)
by: He, Yuchen, et al.
Published: (2024)
Human detectors are surprisingly powerful reward models
by: Ashutosh, Kumar, et al.
Published: (2026)
by: Ashutosh, Kumar, et al.
Published: (2026)
SPOC: Spatially-Progressing Object State Change Segmentation in Video
by: Mandikal, Priyanka, et al.
Published: (2025)
by: Mandikal, Priyanka, et al.
Published: (2025)
Seeing the Arrow of Time in Large Multimodal Models
by: Xue, Zihui, et al.
Published: (2025)
by: Xue, Zihui, et al.
Published: (2025)
Deep Learning for Sports Video Event Detection: Tasks, Datasets, Methods, and Challenges
by: Xu, Hao, et al.
Published: (2025)
by: Xu, Hao, et al.
Published: (2025)
Biomechanical-phase based Temporal Segmentation in Sports Videos: a Demonstration on Javelin-Throw
by: Badatya, Bikash Kumar, et al.
Published: (2025)
by: Badatya, Bikash Kumar, et al.
Published: (2025)
When Thinking Drifts: Evidential Grounding for Robust Video Reasoning
by: Luo, Mi, et al.
Published: (2025)
by: Luo, Mi, et al.
Published: (2025)
DeepSport: A Multimodal Large Language Model for Comprehensive Sports Video Reasoning via Agentic Reinforcement Learning
by: Zou, Junbo, et al.
Published: (2025)
by: Zou, Junbo, et al.
Published: (2025)
ProSkill: Segment-Level Skill Assessment in Procedural Videos
by: Mazzamuto, Michele, et al.
Published: (2026)
by: Mazzamuto, Michele, et al.
Published: (2026)
Towards Temporal Compositional Reasoning in Long-Form Sports Videos
by: Cao, Siyu, et al.
Published: (2026)
by: Cao, Siyu, et al.
Published: (2026)
A General Framework for Jersey Number Recognition in Sports Video
by: Koshkina, Maria, et al.
Published: (2024)
by: Koshkina, Maria, et al.
Published: (2024)
Investigating Event-Based Cameras for Video Frame Interpolation in Sports
by: Deckyvere, Antoine, et al.
Published: (2024)
by: Deckyvere, Antoine, et al.
Published: (2024)
SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports
by: Xia, Haotian, et al.
Published: (2025)
by: Xia, Haotian, et al.
Published: (2025)
Human-in-the-loop Adaptation in Group Activity Feature Learning for Team Sports Video Retrieval
by: Nakatani, Chihiro, et al.
Published: (2026)
by: Nakatani, Chihiro, et al.
Published: (2026)
Similar Items
-
SkillSight: Efficient First-Person Skill Assessment with Gaze
by: Wu, Chi Hsuan, et al.
Published: (2025) -
Learning Skill-Attributes for Transferable Assessment in Video
by: Ashutosh, Kumar, et al.
Published: (2025) -
Stitch-a-Demo: Video Demonstrations from Multistep Descriptions
by: Wu, Chi Hsuan, et al.
Published: (2025) -
ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos
by: Somayazulu, Arjun, et al.
Published: (2026) -
Detours for Navigating Instructional Videos
by: Ashutosh, Kumar, et al.
Published: (2024)