SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Elsharkawi, Ismael, Sait, Ahmed, Giancola, Silvio, Ghanem, Bernard, Sharara, Hossam, Eldesokey, Abdelrahman |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ViG-LRGC: Vision Graph Neural Networks with Learnable Reparameterized Graph Construction
par: Elsharkawi, Ismael, et autres
Publié: (2025)
par: Elsharkawi, Ismael, et autres
Publié: (2025)
Action Anticipation from SoccerNet Football Video Broadcasts
par: Dalal, Mohamad, et autres
Publié: (2025)
par: Dalal, Mohamad, et autres
Publié: (2025)
Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders
par: Eymaël, Alexandre, et autres
Publié: (2024)
par: Eymaël, Alexandre, et autres
Publié: (2024)
SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding
par: Gautam, Sushant, et autres
Publié: (2025)
par: Gautam, Sushant, et autres
Publié: (2025)
SoccerRef-Agents: Multi-Agent System for Automated Soccer Refereeing
par: Meng, Zi, et autres
Publié: (2026)
par: Meng, Zi, et autres
Publié: (2026)
SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries
par: Mkhallati, Hassan, et autres
Publié: (2023)
par: Mkhallati, Hassan, et autres
Publié: (2023)
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
par: Han, Yudong, et autres
Publié: (2024)
par: Han, Yudong, et autres
Publié: (2024)
A Simple Baseline for Streaming Video Understanding
par: Shen, Yujiao, et autres
Publié: (2026)
par: Shen, Yujiao, et autres
Publié: (2026)
Enhancing Road Crack Detection Accuracy with BsS-YOLO: Optimizing Feature Fusion and Attention Mechanisms
par: Tang, Jiaze, et autres
Publié: (2024)
par: Tang, Jiaze, et autres
Publié: (2024)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
par: Li, Huibin, et autres
Publié: (2025)
par: Li, Huibin, et autres
Publié: (2025)
4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding
par: Chen, Zhangquan, et autres
Publié: (2026)
par: Chen, Zhangquan, et autres
Publié: (2026)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
par: Raoufi, Behnam, et autres
Publié: (2025)
par: Raoufi, Behnam, et autres
Publié: (2025)
NAC-TCN: Temporal Convolutional Networks with Causal Dilated Neighborhood Attention for Emotion Understanding
par: Mehta, Alexander, et autres
Publié: (2023)
par: Mehta, Alexander, et autres
Publié: (2023)
Computer Vision for Clinical Gait Analysis: A Gait Abnormality Video Dataset
par: Ranjan, Rahm, et autres
Publié: (2024)
par: Ranjan, Rahm, et autres
Publié: (2024)
VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
par: He, Jianxiang, et autres
Publié: (2025)
par: He, Jianxiang, et autres
Publié: (2025)
Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
par: Li, Yayuan, et autres
Publié: (2025)
par: Li, Yayuan, et autres
Publié: (2025)
SoccerNet-Tracking: Multiple Object Tracking Dataset and Benchmark in Soccer Videos
par: Cioppa, Anthony, et autres
Publié: (2022)
par: Cioppa, Anthony, et autres
Publié: (2022)
Classifying Simulated Gait Impairments using Privacy-preserving Explainable Artificial Intelligence and Mobile Phone Videos
par: Reddy, Lauhitya, et autres
Publié: (2024)
par: Reddy, Lauhitya, et autres
Publié: (2024)
CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models
par: Foss, Aaron, et autres
Publié: (2025)
par: Foss, Aaron, et autres
Publié: (2025)
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
par: Oliveira, Daniel A. P., et autres
Publié: (2025)
par: Oliveira, Daniel A. P., et autres
Publié: (2025)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
par: Anh, Duy Le Dinh, et autres
Publié: (2024)
par: Anh, Duy Le Dinh, et autres
Publié: (2024)
PhysVideoGenerator: Towards Physically Aware Video Generation via Latent Physics Guidance
par: Satish, Siddarth Nilol Kundur, et autres
Publié: (2026)
par: Satish, Siddarth Nilol Kundur, et autres
Publié: (2026)
BlindSight: Harnessing Sparsity for Efficient Vision-Language Models
par: Srikrishnan, Tharun Adithya, et autres
Publié: (2025)
par: Srikrishnan, Tharun Adithya, et autres
Publié: (2025)
Cross-Modal Transfer from Memes to Videos: Addressing Data Scarcity in Hateful Video Detection
par: Wang, Han, et autres
Publié: (2025)
par: Wang, Han, et autres
Publié: (2025)
Beyond Visual Understanding: Introducing PARROT-360V for Vision Language Model Benchmarking
par: Khurdula, Harsha Vardhan, et autres
Publié: (2024)
par: Khurdula, Harsha Vardhan, et autres
Publié: (2024)
WoundNet-Ensemble: A Novel IoMT System Integrating Self-Supervised Deep Learning and Multi-Model Fusion for Automated, High-Accuracy Wound Classification and Healing Progression Monitoring
par: Kiprono, Moses
Publié: (2025)
par: Kiprono, Moses
Publié: (2025)
Web-Scale Collection of Video Data for 4D Animal Reconstruction
par: Zhao, Brian Nlong, et autres
Publié: (2025)
par: Zhao, Brian Nlong, et autres
Publié: (2025)
RDPO: Real Data Preference Optimization for Physics Consistency Video Generation
par: Qian, Wenxu, et autres
Publié: (2025)
par: Qian, Wenxu, et autres
Publié: (2025)
Combined Hyperbolic and Euclidean Soft Triple Loss Beyond the Single Space Deep Metric Learning
par: Saeki, Shozo, et autres
Publié: (2025)
par: Saeki, Shozo, et autres
Publié: (2025)
LESV: Language Embedded Sparse Voxel Fusion for Open-Vocabulary 3D Scene Understanding
par: Wang, Fusang, et autres
Publié: (2026)
par: Wang, Fusang, et autres
Publié: (2026)
GroundCap: A Visually Grounded Image Captioning Dataset
par: Oliveira, Daniel A. P., et autres
Publié: (2025)
par: Oliveira, Daniel A. P., et autres
Publié: (2025)
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
par: Lee, Kyuho, et autres
Publié: (2025)
par: Lee, Kyuho, et autres
Publié: (2025)
VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning
par: Cui, Shaoyang, et autres
Publié: (2026)
par: Cui, Shaoyang, et autres
Publié: (2026)
Context-Aware Network Based on Multi-scale Spatio-temporal Attention for Action Recognition in Videos
par: Li, Xiaoyang, et autres
Publié: (2025)
par: Li, Xiaoyang, et autres
Publié: (2025)
PoseRefer: Pathway-Local Parameters for Semantically Grounded Reference Resolution
par: Deichler, Anna
Publié: (2026)
par: Deichler, Anna
Publié: (2026)
Video Event Reasoning and Prediction by Fusing World Knowledge from LLMs with Vision Foundation Models
par: Dubois, L'ea, et autres
Publié: (2025)
par: Dubois, L'ea, et autres
Publié: (2025)
Boundary-Protection W8A8 HiFloat8 Quantization for Large-Scale Text-to-Video Diffusion Transformers
par: Zhao, Yiming
Publié: (2026)
par: Zhao, Yiming
Publié: (2026)
Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs
par: Feng, Yigui, et autres
Publié: (2026)
par: Feng, Yigui, et autres
Publié: (2026)
Video-CoE: Reinforcing Video Event Prediction via Chain of Events
par: Su, Qile, et autres
Publié: (2026)
par: Su, Qile, et autres
Publié: (2026)
Learning Association via Track-Detection Matching for Multi-Object Tracking
par: Adžemović, Momir
Publié: (2025)
par: Adžemović, Momir
Publié: (2025)
Documents similaires
-
ViG-LRGC: Vision Graph Neural Networks with Learnable Reparameterized Graph Construction
par: Elsharkawi, Ismael, et autres
Publié: (2025) -
Action Anticipation from SoccerNet Football Video Broadcasts
par: Dalal, Mohamad, et autres
Publié: (2025) -
Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders
par: Eymaël, Alexandre, et autres
Publié: (2024) -
SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding
par: Gautam, Sushant, et autres
Publié: (2025) -
SoccerRef-Agents: Multi-Agent System for Automated Soccer Refereeing
par: Meng, Zi, et autres
Publié: (2026)