Learning Consistent Temporal Grounding between Related Tasks in Sports Coaching
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rai, Arushi, Kovashka, Adriana |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generalizing Sports Feedback Generation by Watching Competitions and Reading Books: A Rock Climbing Case Study
von: Rai, Arushi, et al.
Veröffentlicht: (2026)
von: Rai, Arushi, et al.
Veröffentlicht: (2026)
VEIL: Vetting Extracted Image Labels from In-the-Wild Captions for Weakly-Supervised Object Detection
von: Rai, Arushi, et al.
Veröffentlicht: (2023)
von: Rai, Arushi, et al.
Veröffentlicht: (2023)
Role Bias in Diffusion Models: Diagnosing and Mitigating through Intermediate Decomposition
von: Malakouti, Sina, et al.
Veröffentlicht: (2025)
von: Malakouti, Sina, et al.
Veröffentlicht: (2025)
The Face of Persuasion: Analyzing Bias and Generating Culture-Aware Ads
von: Aghazadeh, Aysan, et al.
Veröffentlicht: (2025)
von: Aghazadeh, Aysan, et al.
Veröffentlicht: (2025)
Enhancing Weakly-Supervised Object Detection on Static Images through (Hallucinated) Motion
von: Gungor, Cagri, et al.
Veröffentlicht: (2024)
von: Gungor, Cagri, et al.
Veröffentlicht: (2024)
CAP: Evaluation of Persuasive and Creative Image Generation
von: Aghazadeh, Aysan, et al.
Veröffentlicht: (2024)
von: Aghazadeh, Aysan, et al.
Veröffentlicht: (2024)
Quantifying the Gaps Between Translation and Native Perception in Training for Multimodal, Multilingual Retrieval
von: Buettner, Kyle, et al.
Veröffentlicht: (2024)
von: Buettner, Kyle, et al.
Veröffentlicht: (2024)
Towards Generalization of Tactile Image Generation: Reference-Free Evaluation in a Leakage-Free Setting
von: Gungor, Cagri, et al.
Veröffentlicht: (2025)
von: Gungor, Cagri, et al.
Veröffentlicht: (2025)
Culture in Action: Evaluating Text-to-Image Models through Social Activities
von: Malakouti, Sina, et al.
Veröffentlicht: (2025)
von: Malakouti, Sina, et al.
Veröffentlicht: (2025)
MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding
von: Ran, Ran, et al.
Veröffentlicht: (2026)
von: Ran, Ran, et al.
Veröffentlicht: (2026)
A Multimodal Recaptioning Framework to Account for Perceptual Diversity Across Languages in Vision-Language Modeling
von: Buettner, Kyle, et al.
Veröffentlicht: (2025)
von: Buettner, Kyle, et al.
Veröffentlicht: (2025)
Enhancing Sketch Animation: Text-to-Video Diffusion Models with Temporal Consistency and Rigidity Constraints
von: Rai, Gaurav, et al.
Veröffentlicht: (2024)
von: Rai, Gaurav, et al.
Veröffentlicht: (2024)
Benchmarking VLMs' Reasoning About Persuasive Atypical Images
von: Malakouti, Sina, et al.
Veröffentlicht: (2024)
von: Malakouti, Sina, et al.
Veröffentlicht: (2024)
Leveraging Large Models to Evaluate Novel Content: A Case Study on Advertisement Creativity
von: Hou, Zhaoyi Joey, et al.
Veröffentlicht: (2025)
von: Hou, Zhaoyi Joey, et al.
Veröffentlicht: (2025)
Integrating Audio Narrations to Strengthen Domain Generalization in Multimodal First-Person Action Recognition
von: Gungor, Cagri, et al.
Veröffentlicht: (2024)
von: Gungor, Cagri, et al.
Veröffentlicht: (2024)
Transferring Relative Monocular Depth to Surgical Vision with Temporal Consistency
von: Budd, Charlie, et al.
Veröffentlicht: (2024)
von: Budd, Charlie, et al.
Veröffentlicht: (2024)
From 3D Pose to Prose: Biomechanics-Grounded Vision--Language Coaching
von: Ji, Yuyang, et al.
Veröffentlicht: (2026)
von: Ji, Yuyang, et al.
Veröffentlicht: (2026)
SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs
von: Alansari, Mohamad, et al.
Veröffentlicht: (2026)
von: Alansari, Mohamad, et al.
Veröffentlicht: (2026)
Deep Learning for Sports Video Event Detection: Tasks, Datasets, Methods, and Challenges
von: Xu, Hao, et al.
Veröffentlicht: (2025)
von: Xu, Hao, et al.
Veröffentlicht: (2025)
Towards Understanding Ambiguity Resolution in Multimodal Inference of Meaning
von: Wang, Yufei, et al.
Veröffentlicht: (2025)
von: Wang, Yufei, et al.
Veröffentlicht: (2025)
TechCoach: Towards Technical-Point-Aware Descriptive Action Coaching
von: Li, Yuan-Ming, et al.
Veröffentlicht: (2024)
von: Li, Yuan-Ming, et al.
Veröffentlicht: (2024)
Incorporating Geo-Diverse Knowledge into Prompting for Increased Geographical Robustness in Object Recognition
von: Buettner, Kyle, et al.
Veröffentlicht: (2024)
von: Buettner, Kyle, et al.
Veröffentlicht: (2024)
ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal Guidance
von: Yang, Haijie, et al.
Veröffentlicht: (2024)
von: Yang, Haijie, et al.
Veröffentlicht: (2024)
Vid2Coach: Transforming How-To Videos into Task Assistants
von: Huh, Mina, et al.
Veröffentlicht: (2025)
von: Huh, Mina, et al.
Veröffentlicht: (2025)
STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding
von: Garg, Aaryan, et al.
Veröffentlicht: (2025)
von: Garg, Aaryan, et al.
Veröffentlicht: (2025)
Multi-Scale Contrastive Learning for Video Temporal Grounding
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
Towards Temporal Compositional Reasoning in Long-Form Sports Videos
von: Cao, Siyu, et al.
Veröffentlicht: (2026)
von: Cao, Siyu, et al.
Veröffentlicht: (2026)
Decoupled Spatio-Temporal Consistency Learning for Self-Supervised Tracking
von: Zheng, Yaozong, et al.
Veröffentlicht: (2025)
von: Zheng, Yaozong, et al.
Veröffentlicht: (2025)
VALA: Learning Latent Anchors for Training-Free and Temporally Consistent
von: Wu, Zhangkai, et al.
Veröffentlicht: (2025)
von: Wu, Zhangkai, et al.
Veröffentlicht: (2025)
Temporally Consistent Stereo Matching
von: Zeng, Jiaxi, et al.
Veröffentlicht: (2024)
von: Zeng, Jiaxi, et al.
Veröffentlicht: (2024)
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting
von: Lee, Daeun, et al.
Veröffentlicht: (2026)
von: Lee, Daeun, et al.
Veröffentlicht: (2026)
Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report Generation
von: Yang, Longzhen, et al.
Veröffentlicht: (2025)
von: Yang, Longzhen, et al.
Veröffentlicht: (2025)
Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media Manipulation
von: Li, Yiheng, et al.
Veröffentlicht: (2025)
von: Li, Yiheng, et al.
Veröffentlicht: (2025)
EgoTV: Egocentric Task Verification from Natural Language Task Descriptions
von: Hazra, Rishi, et al.
Veröffentlicht: (2023)
von: Hazra, Rishi, et al.
Veröffentlicht: (2023)
SportSkills: Physical Skill Learning from Sports Instructional Videos
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2026)
von: Ashutosh, Kumar, et al.
Veröffentlicht: (2026)
Multi-Focus Temporal Shifting for Precise Event Spotting in Sports Videos
von: Xu, Hao, et al.
Veröffentlicht: (2025)
von: Xu, Hao, et al.
Veröffentlicht: (2025)
Learning Temporally Consistent Video Depth from Video Diffusion Priors
von: Shao, Jiahao, et al.
Veröffentlicht: (2024)
von: Shao, Jiahao, et al.
Veröffentlicht: (2024)
Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
von: Zheng, Zelin, et al.
Veröffentlicht: (2026)
von: Zheng, Zelin, et al.
Veröffentlicht: (2026)
Temporal Context Consistency Above All: Enhancing Long-Term Anticipation by Learning and Enforcing Temporal Constraints
von: Maté, Alberto, et al.
Veröffentlicht: (2024)
von: Maté, Alberto, et al.
Veröffentlicht: (2024)
GaVS: 3D-Grounded Video Stabilization via Temporally-Consistent Local Reconstruction and Rendering
von: You, Zinuo, et al.
Veröffentlicht: (2025)
von: You, Zinuo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Generalizing Sports Feedback Generation by Watching Competitions and Reading Books: A Rock Climbing Case Study
von: Rai, Arushi, et al.
Veröffentlicht: (2026) -
VEIL: Vetting Extracted Image Labels from In-the-Wild Captions for Weakly-Supervised Object Detection
von: Rai, Arushi, et al.
Veröffentlicht: (2023) -
Role Bias in Diffusion Models: Diagnosing and Mitigating through Intermediate Decomposition
von: Malakouti, Sina, et al.
Veröffentlicht: (2025) -
The Face of Persuasion: Analyzing Bias and Generating Culture-Aware Ads
von: Aghazadeh, Aysan, et al.
Veröffentlicht: (2025) -
Enhancing Weakly-Supervised Object Detection on Static Images through (Hallucinated) Motion
von: Gungor, Cagri, et al.
Veröffentlicht: (2024)