SV3.3B: A Sports Video Understanding Model for Action Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Kodathala, Sai Varun, Vutukoori, Yashwanth Reddy, Vunnam, Rakesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Temporal vs. Spatial: Comparing DINOv3 and V-JEPA2 Feature Representations for Video Action Analysis
by: Kodathala, Sai Varun, et al.
Published: (2025)
by: Kodathala, Sai Varun, et al.
Published: (2025)
The Describe-Then-Generate Bottleneck: How VLM Descriptions Alter Image Generation Outcomes
by: Kodathala, Sai Varun, et al.
Published: (2025)
by: Kodathala, Sai Varun, et al.
Published: (2025)
LLMs can Compress LLMs: Adaptive Pruning by Agents
by: Kodathala, Sai Varun, et al.
Published: (2026)
by: Kodathala, Sai Varun, et al.
Published: (2026)
Can Large Language Models Solve Engineering Equations? A Systematic Comparison of Direct Prediction and Solver-Assisted Approaches
by: Kodathala, Sai Varun, et al.
Published: (2026)
by: Kodathala, Sai Varun, et al.
Published: (2026)
Fast OTSU Thresholding Using Bisection Method
by: Kodathala, Sai Varun
Published: (2025)
by: Kodathala, Sai Varun
Published: (2025)
Six Sigma For Neural Networks: Taguchi-based optimization
by: Kodathala, Sai Varun
Published: (2025)
by: Kodathala, Sai Varun
Published: (2025)
Align before Adapt: Leveraging Entity-to-Region Alignments for Generalizable Video Action Recognition
by: Chen, Yifei, et al.
Published: (2023)
by: Chen, Yifei, et al.
Published: (2023)
Vamos: Versatile Action Models for Video Understanding
by: Wang, Shijie, et al.
Published: (2023)
by: Wang, Shijie, et al.
Published: (2023)
Exploring Explainability in Video Action Recognition
by: Saha, Avinab, et al.
Published: (2024)
by: Saha, Avinab, et al.
Published: (2024)
A Survey on Backbones for Deep Video Action Recognition
by: Tang, Zixuan, et al.
Published: (2024)
by: Tang, Zixuan, et al.
Published: (2024)
SkateboardAI: The Coolest Video Action Recognition for Skateboarding
by: Chen, Hanxiao
Published: (2023)
by: Chen, Hanxiao
Published: (2023)
Flatten: Video Action Recognition is an Image Classification task
by: Chen, Junlin, et al.
Published: (2024)
by: Chen, Junlin, et al.
Published: (2024)
Exploring Ordinal Bias in Action Recognition for Instructional Videos
by: Kim, Joochan, et al.
Published: (2025)
by: Kim, Joochan, et al.
Published: (2025)
Controllable Hybrid Captioner for Improved Long-form Video Understanding
by: Sasse, Kuleen, et al.
Published: (2025)
by: Sasse, Kuleen, et al.
Published: (2025)
DeepSport: A Multimodal Large Language Model for Comprehensive Sports Video Reasoning via Agentic Reinforcement Learning
by: Zou, Junbo, et al.
Published: (2025)
by: Zou, Junbo, et al.
Published: (2025)
InstrAct: Towards Action-Centric Understanding in Instructional Videos
by: Yang, Zhuoyi, et al.
Published: (2026)
by: Yang, Zhuoyi, et al.
Published: (2026)
Temporal Alignment-Free Video Matching for Few-shot Action Recognition
by: Lee, SuBeen, et al.
Published: (2025)
by: Lee, SuBeen, et al.
Published: (2025)
Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using Superquadrics
by: Tse, Tze Ho Elden, et al.
Published: (2025)
by: Tse, Tze Ho Elden, et al.
Published: (2025)
SCBench: A Sports Commentary Benchmark for Video LLMs
by: Ge, Kuangzhi, et al.
Published: (2024)
by: Ge, Kuangzhi, et al.
Published: (2024)
Fire on Motion: Optimizing Video Pass-bands for Efficient Spiking Action Recognition
by: Ye, Shuhan, et al.
Published: (2026)
by: Ye, Shuhan, et al.
Published: (2026)
Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding
by: Mago, Gowreesh, et al.
Published: (2025)
by: Mago, Gowreesh, et al.
Published: (2025)
Efficient Spatial-Temporal Modeling for Real-Time Video Analysis: A Unified Framework for Action Recognition and Object Tracking
by: John, Shahla
Published: (2025)
by: John, Shahla
Published: (2025)
Interpretable Action Recognition on Hard to Classify Actions
by: Anichenko, Anastasia, et al.
Published: (2024)
by: Anichenko, Anastasia, et al.
Published: (2024)
Exploring Audio Hallucination in Egocentric Video Understanding
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
by: Liu, Wenqi, et al.
Published: (2026)
by: Liu, Wenqi, et al.
Published: (2026)
MUSTAN: Multi-scale Temporal Context as Attention for Robust Video Foreground Segmentation
by: Pokala, Praveen Kumar, et al.
Published: (2024)
by: Pokala, Praveen Kumar, et al.
Published: (2024)
S3T-Former: A Purely Spike-Driven State-Space Topology Transformer for Skeleton Action Recognition
by: Zheng, Naichuan, et al.
Published: (2026)
by: Zheng, Naichuan, et al.
Published: (2026)
EPAM-Net: An Efficient Pose-driven Attention-guided Multimodal Network for Video Action Recognition
by: Abdelkawy, Ahmed, et al.
Published: (2024)
by: Abdelkawy, Ahmed, et al.
Published: (2024)
Semi-Supervised Audio-Visual Video Action Recognition with Audio Source Localization Guided Mixup
by: Kang, Seokun, et al.
Published: (2025)
by: Kang, Seokun, et al.
Published: (2025)
A Framework Combining 3D CNN and Transformer for Video-Based Behavior Recognition
by: Zhang, Xiuliang, et al.
Published: (2025)
by: Zhang, Xiuliang, et al.
Published: (2025)
Conformal Predictions for Human Action Recognition with Vision-Language Models
by: Tim, Bary, et al.
Published: (2025)
by: Tim, Bary, et al.
Published: (2025)
3D-RFT: Reinforcement Fine-Tuning for Video-based 3D Scene Understanding
by: Linghu, Xiongkun, et al.
Published: (2026)
by: Linghu, Xiongkun, et al.
Published: (2026)
Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement
by: Li, Bing, et al.
Published: (2022)
by: Li, Bing, et al.
Published: (2022)
Causality Model for Semantic Understanding on Videos
by: Yicong, Li
Published: (2025)
by: Yicong, Li
Published: (2025)
Enhancing Action Recognition by Leveraging the Hierarchical Structure of Actions and Textual Context
by: Benavent-Lledo, Manuel, et al.
Published: (2024)
by: Benavent-Lledo, Manuel, et al.
Published: (2024)
Dual-Signal Adaptive KV-Cache Optimization for Long-Form Video Understanding in Vision-Language Models
by: Sai, Vishnu, et al.
Published: (2026)
by: Sai, Vishnu, et al.
Published: (2026)
MAVIN: Multi-Action Video Generation with Diffusion Models via Transition Video Infilling
by: Zhang, Bowen, et al.
Published: (2024)
by: Zhang, Bowen, et al.
Published: (2024)
FedOnco-Bench: A Reproducible Benchmark for Privacy-Aware Federated Tumor Segmentation with Synthetic CT Data
by: Marella, Viswa Chaitanya, et al.
Published: (2025)
by: Marella, Viswa Chaitanya, et al.
Published: (2025)
Real-Time Drowsiness Detection Using Eye Aspect Ratio and Facial Landmark Detection
by: Rupani, Varun Shiva Krishna, et al.
Published: (2024)
by: Rupani, Varun Shiva Krishna, et al.
Published: (2024)
Efficient Egocentric Action Recognition with Multimodal Data
by: Calzavara, Marco, et al.
Published: (2025)
by: Calzavara, Marco, et al.
Published: (2025)
Similar Items
-
Temporal vs. Spatial: Comparing DINOv3 and V-JEPA2 Feature Representations for Video Action Analysis
by: Kodathala, Sai Varun, et al.
Published: (2025) -
The Describe-Then-Generate Bottleneck: How VLM Descriptions Alter Image Generation Outcomes
by: Kodathala, Sai Varun, et al.
Published: (2025) -
LLMs can Compress LLMs: Adaptive Pruning by Agents
by: Kodathala, Sai Varun, et al.
Published: (2026) -
Can Large Language Models Solve Engineering Equations? A Systematic Comparison of Direct Prediction and Solver-Assisted Approaches
by: Kodathala, Sai Varun, et al.
Published: (2026) -
Fast OTSU Thresholding Using Bisection Method
by: Kodathala, Sai Varun
Published: (2025)