Domain Adaptation of VLM for Soccer Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Tiancheng, Wang, Henry, Salekin, Md Sirajus, Atighehchian, Parmida, Zhang, Shinan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Prompt to Production:Automating Brand-Safe Marketing Imagery with Text-to-Image Models
by: Atighehchian, Parmida, et al.
Published: (2026)
by: Atighehchian, Parmida, et al.
Published: (2026)
DocSplit: A Comprehensive Benchmark Dataset and Evaluation Approach for Document Packet Recognition and Splitting
by: Islam, Md Mofijul, et al.
Published: (2026)
by: Islam, Md Mofijul, et al.
Published: (2026)
SoccerMaster: A Vision Foundation Model for Soccer Understanding
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
AdCare-VLM: Towards a Unified and Pre-aligned Latent Representation for Healthcare Video Understanding
by: Jabin, Md Asaduzzaman, et al.
Published: (2025)
by: Jabin, Md Asaduzzaman, et al.
Published: (2025)
EgoVLM: Policy Optimization for Egocentric Video Understanding
by: Vinod, Ashwin, et al.
Published: (2025)
by: Vinod, Ashwin, et al.
Published: (2025)
SoccerHigh: A Benchmark Dataset for Automatic Soccer Video Summarization
by: Díaz-Juan, Artur, et al.
Published: (2025)
by: Díaz-Juan, Artur, et al.
Published: (2025)
See, Act, Adapt: Active Perception for Unsupervised Cross-Domain Visual Adaptation via Personalized VLM-Guided Agent
by: Tang, Tianci, et al.
Published: (2026)
by: Tang, Tianci, et al.
Published: (2026)
StreamingVLM: Real-Time Understanding for Infinite Video Streams
by: Xu, Ruyi, et al.
Published: (2025)
by: Xu, Ruyi, et al.
Published: (2025)
Agentic generative AI for media content discovery at the national football league
by: Wang, Henry, et al.
Published: (2025)
by: Wang, Henry, et al.
Published: (2025)
VLM6D: VLM based 6Dof Pose Estimation based on RGB-D Images
by: Sarowar, Md Selim, et al.
Published: (2025)
by: Sarowar, Md Selim, et al.
Published: (2025)
RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
by: Liu, Hanqing, et al.
Published: (2026)
by: Liu, Hanqing, et al.
Published: (2026)
Interactive Video Generation via Domain Adaptation
by: Rawal, Ishaan, et al.
Published: (2025)
by: Rawal, Ishaan, et al.
Published: (2025)
AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding
by: Suglia, Alessandro, et al.
Published: (2024)
by: Suglia, Alessandro, et al.
Published: (2024)
AutoSoccerPose: Automated 3D posture Analysis of Soccer Shot Movements
by: Yeung, Calvin, et al.
Published: (2024)
by: Yeung, Calvin, et al.
Published: (2024)
SoccerNet 2023 Challenges Results
by: Cioppa, Anthony, et al.
Published: (2023)
by: Cioppa, Anthony, et al.
Published: (2023)
Towards Universal Soccer Video Understanding
by: Rao, Jiayuan, et al.
Published: (2024)
by: Rao, Jiayuan, et al.
Published: (2024)
Weak to Strong: VLM-Based Pseudo-Labeling as a Weakly Supervised Training Strategy in Multimodal Video-based Hidden Emotion Understanding Tasks
by: Wang, Yufei, et al.
Published: (2026)
by: Wang, Yufei, et al.
Published: (2026)
Modality-Collaborative Low-Rank Decomposers for Few-Shot Video Domain Adaptation
by: Wanyan, Yuyang, et al.
Published: (2025)
by: Wanyan, Yuyang, et al.
Published: (2025)
H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving
by: Chen, Siran, et al.
Published: (2025)
by: Chen, Siran, et al.
Published: (2025)
Bridging Weakly-Supervised Learning and VLM Distillation: Noisy Partial Label Learning for Efficient Downstream Adaptation
by: Wang, Qian-Wei, et al.
Published: (2025)
by: Wang, Qian-Wei, et al.
Published: (2025)
AI Driven Soccer Analysis Using Computer Vision
by: Manchado, Adrian, et al.
Published: (2026)
by: Manchado, Adrian, et al.
Published: (2026)
Unsupervised Domain Adaptation for Action Recognition via Self-Ensembling and Conditional Embedding Alignment
by: Ghosh, Indrajeet, et al.
Published: (2024)
by: Ghosh, Indrajeet, et al.
Published: (2024)
Understanding and Defending VLM Jailbreaks via Jailbreak-Related Representation Shift
by: Wei, Zhihua, et al.
Published: (2026)
by: Wei, Zhihua, et al.
Published: (2026)
City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning
by: Sun, Penglei, et al.
Published: (2025)
by: Sun, Penglei, et al.
Published: (2025)
Rethinking Unsupervised Domain Adaptation for Semantic Segmentation
by: Wang, Zhijie, et al.
Published: (2022)
by: Wang, Zhijie, et al.
Published: (2022)
SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
by: Chen, Pingyi, et al.
Published: (2025)
by: Chen, Pingyi, et al.
Published: (2025)
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
by: Pani, Anupam, et al.
Published: (2025)
by: Pani, Anupam, et al.
Published: (2025)
GazeVLM: A Vision-Language Model for Multi-Task Gaze Understanding
by: Mathew, Athul M., et al.
Published: (2025)
by: Mathew, Athul M., et al.
Published: (2025)
BehaviorVLM: Unified Finetuning-Free Behavioral Understanding with Vision-Language Reasoning
by: Ke, Jingyang, et al.
Published: (2026)
by: Ke, Jingyang, et al.
Published: (2026)
IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model
by: Ji, Yatai, et al.
Published: (2024)
by: Ji, Yatai, et al.
Published: (2024)
Video-As-Prompt: Unified Semantic Control for Video Generation
by: Bian, Yuxuan, et al.
Published: (2025)
by: Bian, Yuxuan, et al.
Published: (2025)
Benchmarking Pathology Foundation Models for Spatial Domain Understanding
by: Zhao, Bokai, et al.
Published: (2026)
by: Zhao, Bokai, et al.
Published: (2026)
VFM-VLM: Vision Foundation Model and Vision Language Model based Visual Comparison for 3D Pose Estimation
by: Sarowar, Md Selim, et al.
Published: (2025)
by: Sarowar, Md Selim, et al.
Published: (2025)
Addressing Domain Shift via Imbalance-Aware Domain Adaptation in Embryo Development Assessment
by: Li, Lei, et al.
Published: (2025)
by: Li, Lei, et al.
Published: (2025)
FOOTPASS: A Multi-Modal Multi-Agent Tactical Context Dataset for Play-by-Play Action Spotting in Soccer Broadcast Videos
by: Ochin, Jeremie, et al.
Published: (2025)
by: Ochin, Jeremie, et al.
Published: (2025)
Feature Fusion Transferability Aware Transformer for Unsupervised Domain Adaptation
by: Yu, Xiaowei, et al.
Published: (2024)
by: Yu, Xiaowei, et al.
Published: (2024)
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
by: Liu, Wenqi, et al.
Published: (2026)
by: Liu, Wenqi, et al.
Published: (2026)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
by: Wang, Shihao, et al.
Published: (2025)
by: Wang, Shihao, et al.
Published: (2025)
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
by: Cai, Yuxuan, et al.
Published: (2025)
by: Cai, Yuxuan, et al.
Published: (2025)
QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design
by: Schneider, Benjamin, et al.
Published: (2025)
by: Schneider, Benjamin, et al.
Published: (2025)
Similar Items
-
From Prompt to Production:Automating Brand-Safe Marketing Imagery with Text-to-Image Models
by: Atighehchian, Parmida, et al.
Published: (2026) -
DocSplit: A Comprehensive Benchmark Dataset and Evaluation Approach for Document Packet Recognition and Splitting
by: Islam, Md Mofijul, et al.
Published: (2026) -
SoccerMaster: A Vision Foundation Model for Soccer Understanding
by: Yang, Haolin, et al.
Published: (2025) -
AdCare-VLM: Towards a Unified and Pre-aligned Latent Representation for Healthcare Video Understanding
by: Jabin, Md Asaduzzaman, et al.
Published: (2025) -
EgoVLM: Policy Optimization for Egocentric Video Understanding
by: Vinod, Ashwin, et al.
Published: (2025)