SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Gautam, Sushant, Midoglu, Cise, Thambawita, Vajira, Riegler, Michael A., Halvorsen, Pål, Shah, Mubarak |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations
di: Gautam, Sushant, et al.
Pubblicazione: (2026)
di: Gautam, Sushant, et al.
Pubblicazione: (2026)
Medico 2025: Visual Question Answering for Gastrointestinal Imaging
di: Gautam, Sushant, et al.
Pubblicazione: (2025)
di: Gautam, Sushant, et al.
Pubblicazione: (2025)
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
di: Gautam, Sushant, et al.
Pubblicazione: (2025)
di: Gautam, Sushant, et al.
Pubblicazione: (2025)
HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
di: Gautam, Sushant, et al.
Pubblicazione: (2025)
di: Gautam, Sushant, et al.
Pubblicazione: (2025)
SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset
di: Gautam, Sushant, et al.
Pubblicazione: (2024)
di: Gautam, Sushant, et al.
Pubblicazione: (2024)
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
di: Zhang, Junwen, et al.
Pubblicazione: (2025)
di: Zhang, Junwen, et al.
Pubblicazione: (2025)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
di: Viveiros, André G., et al.
Pubblicazione: (2025)
di: Viveiros, André G., et al.
Pubblicazione: (2025)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
di: Yim, Wen-wai, et al.
Pubblicazione: (2025)
di: Yim, Wen-wai, et al.
Pubblicazione: (2025)
Unpacking Hateful Memes: Presupposed Context and False Claims
di: Cai, Weibin, et al.
Pubblicazione: (2025)
di: Cai, Weibin, et al.
Pubblicazione: (2025)
Key-Scan-Based Mobile Robot Navigation: Integrated Mapping, Planning, and Control using Graphs of Scan Regions
di: Latha, Dharshan Bashkaran, et al.
Pubblicazione: (2024)
di: Latha, Dharshan Bashkaran, et al.
Pubblicazione: (2024)
Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy
di: Gautam, Sushant, et al.
Pubblicazione: (2025)
di: Gautam, Sushant, et al.
Pubblicazione: (2025)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
di: Huo, Dongjie, et al.
Pubblicazione: (2026)
di: Huo, Dongjie, et al.
Pubblicazione: (2026)
Floorplan2Guide: LLM-Guided Floorplan Parsing for BLV Indoor Navigation
di: Ayanzadeh, Aydin, et al.
Pubblicazione: (2025)
di: Ayanzadeh, Aydin, et al.
Pubblicazione: (2025)
Spectral Integrated Gradients for Coarse-to-Fine Feature Attribution
di: Kim, Soyeon, et al.
Pubblicazione: (2026)
di: Kim, Soyeon, et al.
Pubblicazione: (2026)
Manifold-Aligned Guided Integrated Gradients for Reliable Feature Attribution
di: Kim, Soyeon, et al.
Pubblicazione: (2026)
di: Kim, Soyeon, et al.
Pubblicazione: (2026)
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
di: Sáez, Arnau Igualde, et al.
Pubblicazione: (2025)
di: Sáez, Arnau Igualde, et al.
Pubblicazione: (2025)
Context-Dependent Affordance Computation in Vision-Language Models
di: Farzulla, Murad
Pubblicazione: (2026)
di: Farzulla, Murad
Pubblicazione: (2026)
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
di: Fu, Tianyu, et al.
Pubblicazione: (2024)
di: Fu, Tianyu, et al.
Pubblicazione: (2024)
The MSR-Video to Text Dataset with Clean Annotations
di: Chen, Haoran, et al.
Pubblicazione: (2021)
di: Chen, Haoran, et al.
Pubblicazione: (2021)
SoccerRAG: Multimodal Soccer Information Retrieval via Natural Queries
di: Strand, Aleksander Theo, et al.
Pubblicazione: (2024)
di: Strand, Aleksander Theo, et al.
Pubblicazione: (2024)
MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation
di: Qi, Dekang, et al.
Pubblicazione: (2026)
di: Qi, Dekang, et al.
Pubblicazione: (2026)
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
di: Yanambakkam, Hemanth Teja, et al.
Pubblicazione: (2025)
di: Yanambakkam, Hemanth Teja, et al.
Pubblicazione: (2025)
Extraction Of Cumulative Blobs From Dynamic Gestures
di: Naulakha, Rishabh, et al.
Pubblicazione: (2025)
di: Naulakha, Rishabh, et al.
Pubblicazione: (2025)
Polarization-Based Eye Tracking with Personalized Siamese Architectures
di: Kalkanli, Beyza, et al.
Pubblicazione: (2026)
di: Kalkanli, Beyza, et al.
Pubblicazione: (2026)
A Survey on Vision-Language-Action Models for Embodied AI
di: Ma, Yueen, et al.
Pubblicazione: (2024)
di: Ma, Yueen, et al.
Pubblicazione: (2024)
Surg$Σ$: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence
di: Zeng, Zhitao, et al.
Pubblicazione: (2026)
di: Zeng, Zhitao, et al.
Pubblicazione: (2026)
SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations
di: Dumpala, Sri Harsha, et al.
Pubblicazione: (2024)
di: Dumpala, Sri Harsha, et al.
Pubblicazione: (2024)
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
di: Chen, Kewei, et al.
Pubblicazione: (2025)
di: Chen, Kewei, et al.
Pubblicazione: (2025)
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
di: Chen, Kewei, et al.
Pubblicazione: (2026)
di: Chen, Kewei, et al.
Pubblicazione: (2026)
Does CLIP perceive art the same way we do?
di: Asperti, Andrea, et al.
Pubblicazione: (2025)
di: Asperti, Andrea, et al.
Pubblicazione: (2025)
Demo: Soccer Information Retrieval via Natural Queries using SoccerRAG
di: Strand, Aleksander Theo, et al.
Pubblicazione: (2024)
di: Strand, Aleksander Theo, et al.
Pubblicazione: (2024)
Multimodal Structure-Aware Quantum Data Processing
di: Hawashin, Hala, et al.
Pubblicazione: (2024)
di: Hawashin, Hala, et al.
Pubblicazione: (2024)
VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
di: Yang, Baoyao, et al.
Pubblicazione: (2025)
di: Yang, Baoyao, et al.
Pubblicazione: (2025)
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
di: Tu, Songjun, et al.
Pubblicazione: (2025)
di: Tu, Songjun, et al.
Pubblicazione: (2025)
Tricks and Plug-ins for Gradient Boosting in Image Classification
di: Fang, Biyi, et al.
Pubblicazione: (2025)
di: Fang, Biyi, et al.
Pubblicazione: (2025)
Clarification as Supervision: Reinforcement Learning for Vision-Language Interfaces
di: Gkountouras, John, et al.
Pubblicazione: (2025)
di: Gkountouras, John, et al.
Pubblicazione: (2025)
RACAS: Controlling Diverse Robots With a Single Agentic System
di: Ashley, Dylan R., et al.
Pubblicazione: (2026)
di: Ashley, Dylan R., et al.
Pubblicazione: (2026)
CulinaryCut-VLAP: A Vision-Language-Action-Physics Framework for Food Cutting via a Force-Aware Material Point Method
di: Koh, Hyunseo, et al.
Pubblicazione: (2026)
di: Koh, Hyunseo, et al.
Pubblicazione: (2026)
AI-based Drone Assisted Human Rescue in Disaster Environments: Challenges and Opportunities
di: Papyan, Narek, et al.
Pubblicazione: (2024)
di: Papyan, Narek, et al.
Pubblicazione: (2024)
ReasonPlan: Unified Scene Prediction and Decision Reasoning for Closed-loop Autonomous Driving
di: Liu, Xueyi, et al.
Pubblicazione: (2025)
di: Liu, Xueyi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations
di: Gautam, Sushant, et al.
Pubblicazione: (2026) -
Medico 2025: Visual Question Answering for Gastrointestinal Imaging
di: Gautam, Sushant, et al.
Pubblicazione: (2025) -
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
di: Gautam, Sushant, et al.
Pubblicazione: (2025) -
HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
di: Gautam, Sushant, et al.
Pubblicazione: (2025) -
SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset
di: Gautam, Sushant, et al.
Pubblicazione: (2024)