Do We Need Large VLMs for Spotting Soccer Actions?
Fuente:
arXiv
Salvato in:
| Autori principali: | Chakraborty, Ritabrata, Chakraborty, Rajatsubhra, Dasgupta, Avijit, Chaurasia, Sandeep |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CAMBench-QR : A Structure-Aware Benchmark for Post-Hoc Explanations with QR Understanding
di: Chakraborty, Ritabrata, et al.
Pubblicazione: (2025)
di: Chakraborty, Ritabrata, et al.
Pubblicazione: (2025)
TruthLens:A Training-Free Paradigm for DeepFake Detection
di: Chakraborty, Ritabrata, et al.
Pubblicazione: (2025)
di: Chakraborty, Ritabrata, et al.
Pubblicazione: (2025)
Survey of Action Recognition, Spotting and Spatio-Temporal Localization in Soccer -- Current Trends and Research Perspectives
di: Seweryn, Karolina, et al.
Pubblicazione: (2023)
di: Seweryn, Karolina, et al.
Pubblicazione: (2023)
MambaOut: Do We Really Need Mamba for Vision?
di: Yu, Weihao, et al.
Pubblicazione: (2024)
di: Yu, Weihao, et al.
Pubblicazione: (2024)
Mask & Match: Learning to Recognize Handwritten Math with Self-Supervised Attention
di: Mitra, Shree, et al.
Pubblicazione: (2025)
di: Mitra, Shree, et al.
Pubblicazione: (2025)
Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
di: Li, Kevin Y., et al.
Pubblicazione: (2024)
di: Li, Kevin Y., et al.
Pubblicazione: (2024)
DM-QPMNET: Dual-modality fusion network for cell segmentation in quantitative phase microscopy
di: Chakraborty, Rajatsubhra, et al.
Pubblicazione: (2025)
di: Chakraborty, Rajatsubhra, et al.
Pubblicazione: (2025)
Towards Precise Action Spotting: Addressing Temporal Misalignment in Labels with Dynamic Label Assignment
di: Tamura, Masato
Pubblicazione: (2025)
di: Tamura, Masato
Pubblicazione: (2025)
Do We Really Need Quantum Machine Learning?: A Multidimensional Empirical Study
di: Vhaduri, Sudip, et al.
Pubblicazione: (2026)
di: Vhaduri, Sudip, et al.
Pubblicazione: (2026)
Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
di: Wang, Yanbo, et al.
Pubblicazione: (2025)
di: Wang, Yanbo, et al.
Pubblicazione: (2025)
BiasMap: Leveraging Cross-Attentions to Discover and Mitigate Hidden Social Biases in Text-to-Image Generation
di: Chakraborty, Rajatsubhra, et al.
Pubblicazione: (2025)
di: Chakraborty, Rajatsubhra, et al.
Pubblicazione: (2025)
A Transformer Based Handwriting Recognition System Jointly Using Online and Offline Features
di: Lodh, Ayush, et al.
Pubblicazione: (2025)
di: Lodh, Ayush, et al.
Pubblicazione: (2025)
Towards Robust Cross-Dataset Object Detection Generalization under Domain Specificity
di: Chakraborty, Ritabrata, et al.
Pubblicazione: (2026)
di: Chakraborty, Ritabrata, et al.
Pubblicazione: (2026)
A Lightweight Context-Driven Training-Free Network for Scene Text Segmentation and Recognition
di: Chakraborty, Ritabrata, et al.
Pubblicazione: (2025)
di: Chakraborty, Ritabrata, et al.
Pubblicazione: (2025)
Rethinking VLMs and LLMs for Image Classification
di: Cooper, Avi, et al.
Pubblicazione: (2024)
di: Cooper, Avi, et al.
Pubblicazione: (2024)
Sketch-guided Image Inpainting with Partial Discrete Diffusion Process
di: Sharma, Nakul, et al.
Pubblicazione: (2024)
di: Sharma, Nakul, et al.
Pubblicazione: (2024)
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
di: Li, Shuo, et al.
Pubblicazione: (2024)
di: Li, Shuo, et al.
Pubblicazione: (2024)
CARES: Context-Aware Resolution Selector for VLMs
di: Kimhi, Moshe, et al.
Pubblicazione: (2025)
di: Kimhi, Moshe, et al.
Pubblicazione: (2025)
DASH: Detection and Assessment of Systematic Hallucinations of VLMs
di: Augustin, Maximilian, et al.
Pubblicazione: (2025)
di: Augustin, Maximilian, et al.
Pubblicazione: (2025)
Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis
di: Moon, Jihyun, et al.
Pubblicazione: (2025)
di: Moon, Jihyun, et al.
Pubblicazione: (2025)
Hidden in plain sight: VLMs overlook their visual representations
di: Fu, Stephanie, et al.
Pubblicazione: (2025)
di: Fu, Stephanie, et al.
Pubblicazione: (2025)
Quo Vadis, Anomaly Detection? LLMs and VLMs in the Spotlight
di: Ding, Xi, et al.
Pubblicazione: (2024)
di: Ding, Xi, et al.
Pubblicazione: (2024)
DEAL: Disentangle and Localize Concept-level Explanations for VLMs
di: Li, Tang, et al.
Pubblicazione: (2024)
di: Li, Tang, et al.
Pubblicazione: (2024)
LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living
di: Reilly, Dominick, et al.
Pubblicazione: (2024)
di: Reilly, Dominick, et al.
Pubblicazione: (2024)
SoccerNet Game State Reconstruction: End-to-End Athlete Tracking and Identification on a Minimap
di: Somers, Vladimir, et al.
Pubblicazione: (2024)
di: Somers, Vladimir, et al.
Pubblicazione: (2024)
Galaxy Walker: Geometry-aware VLMs For Galaxy-scale Understanding
di: Chen, Tianyu, et al.
Pubblicazione: (2025)
di: Chen, Tianyu, et al.
Pubblicazione: (2025)
IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
di: Faraz, Ali, et al.
Pubblicazione: (2025)
di: Faraz, Ali, et al.
Pubblicazione: (2025)
ArtifactLens: Hundreds of Labels Are Enough for Artifact Detection with VLMs
di: Burgess, James, et al.
Pubblicazione: (2026)
di: Burgess, James, et al.
Pubblicazione: (2026)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
di: Zhang, Jianrui, et al.
Pubblicazione: (2026)
di: Zhang, Jianrui, et al.
Pubblicazione: (2026)
Can VLMs Reason Robustly? A Neuro-Symbolic Investigation
di: Chen, Weixin, et al.
Pubblicazione: (2026)
di: Chen, Weixin, et al.
Pubblicazione: (2026)
Fast & Efficient Normalizing Flows and Applications of Image Generative Models
di: Nagar, Sandeep
Pubblicazione: (2025)
di: Nagar, Sandeep
Pubblicazione: (2025)
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
di: Ryoo, Michael S., et al.
Pubblicazione: (2024)
di: Ryoo, Michael S., et al.
Pubblicazione: (2024)
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
di: Izadi, Amirmohammad, et al.
Pubblicazione: (2025)
di: Izadi, Amirmohammad, et al.
Pubblicazione: (2025)
COVR:Collaborative Optimization of VLMs and RL Agent for Visual-Based Control
di: Xia, Canming, et al.
Pubblicazione: (2026)
di: Xia, Canming, et al.
Pubblicazione: (2026)
Do We Really Need a Large Number of Visual Prompts?
di: Kim, Youngeun, et al.
Pubblicazione: (2023)
di: Kim, Youngeun, et al.
Pubblicazione: (2023)
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints
di: Kumar, Sunil, et al.
Pubblicazione: (2025)
di: Kumar, Sunil, et al.
Pubblicazione: (2025)
Right this way: Can VLMs Guide Us to See More to Answer Questions?
di: Liu, Li, et al.
Pubblicazione: (2024)
di: Liu, Li, et al.
Pubblicazione: (2024)
Discovering Closed-Loop Failures of Vision-Based Controllers via Reachability Analysis
di: Chakraborty, Kaustav, et al.
Pubblicazione: (2022)
di: Chakraborty, Kaustav, et al.
Pubblicazione: (2022)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
di: Qin, Yiming, et al.
Pubblicazione: (2025)
di: Qin, Yiming, et al.
Pubblicazione: (2025)
An Expert Ensemble for Detecting Anomalous Scenes, Interactions, and Behaviors in Autonomous Driving
di: Ji, Tianchen, et al.
Pubblicazione: (2025)
di: Ji, Tianchen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
CAMBench-QR : A Structure-Aware Benchmark for Post-Hoc Explanations with QR Understanding
di: Chakraborty, Ritabrata, et al.
Pubblicazione: (2025) -
TruthLens:A Training-Free Paradigm for DeepFake Detection
di: Chakraborty, Ritabrata, et al.
Pubblicazione: (2025) -
Survey of Action Recognition, Spotting and Spatio-Temporal Localization in Soccer -- Current Trends and Research Perspectives
di: Seweryn, Karolina, et al.
Pubblicazione: (2023) -
MambaOut: Do We Really Need Mamba for Vision?
di: Yu, Weihao, et al.
Pubblicazione: (2024) -
Mask & Match: Learning to Recognize Handwritten Math with Self-Supervised Attention
di: Mitra, Shree, et al.
Pubblicazione: (2025)