OpenTAD: A Unified Framework and Comprehensive Study of Temporal Action Detection
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Shuming, Zhao, Chen, Zohra, Fatimah, Soldan, Mattia, Pardo, Alejandro, Xu, Mengmeng, Alssum, Lama, Ramazanova, Merey, Alcázar, Juan León, Cioppa, Anthony, Giancola, Silvio, Hinojosa, Carlos, Ghanem, Bernard |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Pixels or Positions? Benchmarking Modalities in Group Activity Recognition
por: Karki, Drishya, et al.
Publicado: (2025)
por: Karki, Drishya, et al.
Publicado: (2025)
GoTTA be Diverse: Rethinking Memory Policies for Test-Time Adaptation
por: Alhuwaider, Shyma, et al.
Publicado: (2026)
por: Alhuwaider, Shyma, et al.
Publicado: (2026)
OSL-ActionSpotting: A Unified Library for Action Spotting in Sports Videos
por: Benzakour, Yassine, et al.
Publicado: (2024)
por: Benzakour, Yassine, et al.
Publicado: (2024)
Exploring Missing Modality in Multimodal Egocentric Datasets
por: Ramazanova, Merey, et al.
Publicado: (2024)
por: Ramazanova, Merey, et al.
Publicado: (2024)
Test-Time Adaptation for Combating Missing Modalities in Egocentric Videos
por: Ramazanova, Merey, et al.
Publicado: (2024)
por: Ramazanova, Merey, et al.
Publicado: (2024)
Deep learning for action spotting in association football videos
por: Giancola, Silvio, et al.
Publicado: (2024)
por: Giancola, Silvio, et al.
Publicado: (2024)
SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries
por: Mkhallati, Hassan, et al.
Publicado: (2023)
por: Mkhallati, Hassan, et al.
Publicado: (2023)
Investigating Event-Based Cameras for Video Frame Interpolation in Sports
por: Deckyvere, Antoine, et al.
Publicado: (2024)
por: Deckyvere, Antoine, et al.
Publicado: (2024)
Learning Semantic Segmentation with Query Points Supervision on Aerial Images
por: Rivier, Santiago, et al.
Publicado: (2023)
por: Rivier, Santiago, et al.
Publicado: (2023)
X-VARS: Introducing Explainability in Football Refereeing with Multi-Modal Large Language Model
por: Held, Jan, et al.
Publicado: (2024)
por: Held, Jan, et al.
Publicado: (2024)
Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders
por: Eymaël, Alexandre, et al.
Publicado: (2024)
por: Eymaël, Alexandre, et al.
Publicado: (2024)
VARS: Video Assistant Referee System for Automated Soccer Decision Making from Multiple Views
por: Held, Jan, et al.
Publicado: (2023)
por: Held, Jan, et al.
Publicado: (2023)
Towards Active Learning for Action Spotting in Association Football Videos
por: Giancola, Silvio, et al.
Publicado: (2023)
por: Giancola, Silvio, et al.
Publicado: (2023)
ADVMEM: Adversarial Memory Initialization for Realistic Test-Time Adaptation via Tracklet-Based Benchmarking
por: Alhuwaider, Shyma, et al.
Publicado: (2025)
por: Alhuwaider, Shyma, et al.
Publicado: (2025)
Towards AI-Powered Video Assistant Referee System (VARS) for Association Football
por: Held, Jan, et al.
Publicado: (2024)
por: Held, Jan, et al.
Publicado: (2024)
Action Anticipation from SoccerNet Football Video Broadcasts
por: Dalal, Mohamad, et al.
Publicado: (2025)
por: Dalal, Mohamad, et al.
Publicado: (2025)
Forget Less, Retain More: A Lightweight Regularizer for Rehearsal-Based Continual Learning
por: Alssum, Lama, et al.
Publicado: (2025)
por: Alssum, Lama, et al.
Publicado: (2025)
ColorMAE: Exploring data-independent masking strategies in Masked AutoEncoders
por: Hinojosa, Carlos, et al.
Publicado: (2024)
por: Hinojosa, Carlos, et al.
Publicado: (2024)
$β$-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment
por: Zohra, Fatimah, et al.
Publicado: (2025)
por: Zohra, Fatimah, et al.
Publicado: (2025)
Compressed-Language Models for Understanding Compressed File Formats: a JPEG Exploration
por: Pérez, Juan C., et al.
Publicado: (2024)
por: Pérez, Juan C., et al.
Publicado: (2024)
SoccerNet-Tracking: Multiple Object Tracking Dataset and Benchmark in Soccer Videos
por: Cioppa, Anthony, et al.
Publicado: (2022)
por: Cioppa, Anthony, et al.
Publicado: (2022)
Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation
por: Eldesokey, Abdelrahman, et al.
Publicado: (2026)
por: Eldesokey, Abdelrahman, et al.
Publicado: (2026)
ResidualViT for Efficient Temporally Dense Video Encoding
por: Soldan, Mattia, et al.
Publicado: (2025)
por: Soldan, Mattia, et al.
Publicado: (2025)
MVTN: Learning Multi-View Transformations for 3D Understanding
por: Hamdi, Abdullah, et al.
Publicado: (2022)
por: Hamdi, Abdullah, et al.
Publicado: (2022)
3D Convex Splatting: Radiance Field Rendering with 3D Smooth Convexes
por: Held, Jan, et al.
Publicado: (2024)
por: Held, Jan, et al.
Publicado: (2024)
Evaluation of Test-Time Adaptation Under Computational Time Constraints
por: Alfarra, Motasem, et al.
Publicado: (2023)
por: Alfarra, Motasem, et al.
Publicado: (2023)
End-to-End Temporal Action Detection with 1B Parameters Across 1000 Frames
por: Liu, Shuming, et al.
Publicado: (2023)
por: Liu, Shuming, et al.
Publicado: (2023)
Dr$^2$Net: Dynamic Reversible Dual-Residual Networks for Memory-Efficient Finetuning
por: Zhao, Chen, et al.
Publicado: (2024)
por: Zhao, Chen, et al.
Publicado: (2024)
Harnessing Temporal Causality for Advanced Temporal Action Detection
por: Liu, Shuming, et al.
Publicado: (2024)
por: Liu, Shuming, et al.
Publicado: (2024)
Transformers from Compressed Representations
por: Alcazar, Juan C. Leon, et al.
Publicado: (2025)
por: Alcazar, Juan C. Leon, et al.
Publicado: (2025)
Hybrid Structure-from-Motion and Camera Relocalization for Enhanced Egocentric Localization
por: Mai, Jinjie, et al.
Publicado: (2024)
por: Mai, Jinjie, et al.
Publicado: (2024)
Triangle Splatting for Real-Time Radiance Field Rendering
por: Held, Jan, et al.
Publicado: (2025)
por: Held, Jan, et al.
Publicado: (2025)
Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
por: Alssum, Lama, et al.
Publicado: (2025)
por: Alssum, Lama, et al.
Publicado: (2025)
SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy
por: Elsharkawi, Ismael, et al.
Publicado: (2026)
por: Elsharkawi, Ismael, et al.
Publicado: (2026)
Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
por: Kim, Minseon, et al.
Publicado: (2025)
por: Kim, Minseon, et al.
Publicado: (2025)
Video Self-Stitching Graph Network for Temporal Action Localization
por: Zhao, Chen, et al.
Publicado: (2020)
por: Zhao, Chen, et al.
Publicado: (2020)
SAVeS: Steering Safety Judgments in Vision-Language Models via Semantic Cues
por: Hinojosa, Carlos, et al.
Publicado: (2026)
por: Hinojosa, Carlos, et al.
Publicado: (2026)
CAPTAIN: Semantic Feature Injection for Memorization Mitigation in Text-to-Image Diffusion Models
por: Zhang, Tong, et al.
Publicado: (2025)
por: Zhang, Tong, et al.
Publicado: (2025)
Towards Automated Movie Trailer Generation
por: Argaw, Dawit Mureja, et al.
Publicado: (2024)
por: Argaw, Dawit Mureja, et al.
Publicado: (2024)
BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding
por: Liu, Shuming, et al.
Publicado: (2025)
por: Liu, Shuming, et al.
Publicado: (2025)
Ejemplares similares
-
Pixels or Positions? Benchmarking Modalities in Group Activity Recognition
por: Karki, Drishya, et al.
Publicado: (2025) -
GoTTA be Diverse: Rethinking Memory Policies for Test-Time Adaptation
por: Alhuwaider, Shyma, et al.
Publicado: (2026) -
OSL-ActionSpotting: A Unified Library for Action Spotting in Sports Videos
por: Benzakour, Yassine, et al.
Publicado: (2024) -
Exploring Missing Modality in Multimodal Egocentric Datasets
por: Ramazanova, Merey, et al.
Publicado: (2024) -
Test-Time Adaptation for Combating Missing Modalities in Egocentric Videos
por: Ramazanova, Merey, et al.
Publicado: (2024)