Towards Universal Soccer Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Rao, Jiayuan, Wu, Haoning, Jiang, Hao, Zhang, Ya, Wang, Yanfeng, Xie, Weidi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Agent System for Comprehensive Soccer Understanding
by: Rao, Jiayuan, et al.
Published: (2025)
by: Rao, Jiayuan, et al.
Published: (2025)
MatchTime: Towards Automatic Soccer Game Commentary Generation
by: Rao, Jiayuan, et al.
Published: (2024)
by: Rao, Jiayuan, et al.
Published: (2024)
SoccerMaster: A Vision Foundation Model for Soccer Understanding
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence
by: Wu, Haoning, et al.
Published: (2025)
by: Wu, Haoning, et al.
Published: (2025)
MRGen: Segmentation Data Engine for Underrepresented MRI Modalities
by: Wu, Haoning, et al.
Published: (2024)
by: Wu, Haoning, et al.
Published: (2024)
SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass
by: Meng, Yanxu, et al.
Published: (2025)
by: Meng, Yanxu, et al.
Published: (2025)
Intelligent Grimm -- Open-ended Visual Storytelling via Latent Diffusion Models
by: Liu, Chang, et al.
Published: (2023)
by: Liu, Chang, et al.
Published: (2023)
Count Anything at Any Granularity
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
Multi-Sentence Grounding for Long-term Instructional Video
by: Li, Zeqian, et al.
Published: (2023)
by: Li, Zeqian, et al.
Published: (2023)
Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
by: Li, Zeqian, et al.
Published: (2025)
by: Li, Zeqian, et al.
Published: (2025)
Zero-shot Composed Text-Image Retrieval
by: Liu, Yikun, et al.
Published: (2023)
by: Liu, Yikun, et al.
Published: (2023)
Knowledge-enhanced Visual-Language Pretraining for Computational Pathology
by: Zhou, Xiao, et al.
Published: (2024)
by: Zhou, Xiao, et al.
Published: (2024)
Deep Understanding of Soccer Match Videos
by: Xu, Shikun, et al.
Published: (2024)
by: Xu, Shikun, et al.
Published: (2024)
PhenoLIP: Integrating Phenotype Ontology Knowledge into Medical Vision-Language Pretraining
by: Liang, Cheng, et al.
Published: (2026)
by: Liang, Cheng, et al.
Published: (2026)
Domain Adaptation of VLM for Soccer Video Understanding
by: Jiang, Tiancheng, et al.
Published: (2025)
by: Jiang, Tiancheng, et al.
Published: (2025)
RadGenome-Chest CT: A Grounded Vision-Language Dataset for Chest CT Analysis
by: Zhang, Xiaoman, et al.
Published: (2024)
by: Zhang, Xiaoman, et al.
Published: (2024)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
by: Zhang, Xiaoman, et al.
Published: (2023)
by: Zhang, Xiaoman, et al.
Published: (2023)
OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams
by: Yan, Yibin, et al.
Published: (2026)
by: Yan, Yibin, et al.
Published: (2026)
How Well Can Modern LLMs Act as Agent Cores in Radiology Environments?
by: Zheng, Qiaoyu, et al.
Published: (2024)
by: Zheng, Qiaoyu, et al.
Published: (2024)
MegaFusion: Extend Diffusion Models towards Higher-resolution Image Generation without Further Tuning
by: Wu, Haoning, et al.
Published: (2024)
by: Wu, Haoning, et al.
Published: (2024)
A Sanity Check on Composed Image Retrieval
by: Liu, Yikun, et al.
Published: (2026)
by: Liu, Yikun, et al.
Published: (2026)
ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification
by: Fan, Ziqing, et al.
Published: (2025)
by: Fan, Ziqing, et al.
Published: (2025)
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023)
by: Di, Shangzhe, et al.
Published: (2023)
GenTac: Generative Modeling and Forecasting of Soccer Tactics
by: Rao, Jiayuan, et al.
Published: (2026)
by: Rao, Jiayuan, et al.
Published: (2026)
Rethinking Whole-Body CT Image Interpretation: An Abnormality-Centric Approach
by: Zhao, Ziheng, et al.
Published: (2025)
by: Zhao, Ziheng, et al.
Published: (2025)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
by: Chen, Qirui, et al.
Published: (2024)
by: Chen, Qirui, et al.
Published: (2024)
SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy
by: Elsharkawi, Ismael, et al.
Published: (2026)
by: Elsharkawi, Ismael, et al.
Published: (2026)
POINTS-Seeker: Towards Training a Multimodal Agentic Search Model from Scratch
by: Liu, Yikun, et al.
Published: (2026)
by: Liu, Yikun, et al.
Published: (2026)
M^3Builder: A Multi-Agent System for Automated Machine Learning in Medical Imaging
by: Feng, Jinghao, et al.
Published: (2025)
by: Feng, Jinghao, et al.
Published: (2025)
LoRKD: Low-Rank Knowledge Decomposition for Medical Foundation Models
by: Li, Haolin, et al.
Published: (2024)
by: Li, Haolin, et al.
Published: (2024)
Large-scale Long-tailed Disease Diagnosis on Radiology Images
by: Zheng, Qiaoyu, et al.
Published: (2023)
by: Zheng, Qiaoyu, et al.
Published: (2023)
Large-Vocabulary Segmentation for Medical Images with Text Prompts
by: Zhao, Ziheng, et al.
Published: (2023)
by: Zhao, Ziheng, et al.
Published: (2023)
SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries
by: Mkhallati, Hassan, et al.
Published: (2023)
by: Mkhallati, Hassan, et al.
Published: (2023)
Character-Centric Understanding of Animated Movies
by: Gui, Zhongrui, et al.
Published: (2025)
by: Gui, Zhongrui, et al.
Published: (2025)
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
by: Wu, Haoning, et al.
Published: (2024)
by: Wu, Haoning, et al.
Published: (2024)
RadIR: A Scalable Framework for Multi-Grained Medical Image Retrieval via Radiology Report Mining
by: Zhang, Tengfei, et al.
Published: (2025)
by: Zhang, Tengfei, et al.
Published: (2025)
SoccerNet-Tracking: Multiple Object Tracking Dataset and Benchmark in Soccer Videos
by: Cioppa, Anthony, et al.
Published: (2022)
by: Cioppa, Anthony, et al.
Published: (2022)
Generative Frame Sampler for Long Video Understanding
by: Yao, Linli, et al.
Published: (2025)
by: Yao, Linli, et al.
Published: (2025)
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs
by: Zhang, Zicheng, et al.
Published: (2024)
by: Zhang, Zicheng, et al.
Published: (2024)
Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation
by: Shi, Yudi, et al.
Published: (2024)
by: Shi, Yudi, et al.
Published: (2024)
Similar Items
-
Multi-Agent System for Comprehensive Soccer Understanding
by: Rao, Jiayuan, et al.
Published: (2025) -
MatchTime: Towards Automatic Soccer Game Commentary Generation
by: Rao, Jiayuan, et al.
Published: (2024) -
SoccerMaster: A Vision Foundation Model for Soccer Understanding
by: Yang, Haolin, et al.
Published: (2025) -
SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence
by: Wu, Haoning, et al.
Published: (2025) -
MRGen: Segmentation Data Engine for Underrepresented MRI Modalities
by: Wu, Haoning, et al.
Published: (2024)