End-to-End Temporal Action Detection with 1B Parameters Across 1000 Frames
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Shuming, Zhang, Chen-Lin, Zhao, Chen, Ghanem, Bernard |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Harnessing Temporal Causality for Advanced Temporal Action Detection
by: Liu, Shuming, et al.
Published: (2024)
by: Liu, Shuming, et al.
Published: (2024)
TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos
by: Zhang, Chen-Lin, et al.
Published: (2025)
by: Zhang, Chen-Lin, et al.
Published: (2025)
Video Self-Stitching Graph Network for Temporal Action Localization
by: Zhao, Chen, et al.
Published: (2020)
by: Zhao, Chen, et al.
Published: (2020)
BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding
by: Liu, Shuming, et al.
Published: (2025)
by: Liu, Shuming, et al.
Published: (2025)
End-to-End Streaming Video Temporal Action Segmentation with Reinforce Learning
by: Zhang, Jinrong, et al.
Published: (2023)
by: Zhang, Jinrong, et al.
Published: (2023)
OpenTAD: A Unified Framework and Comprehensive Study of Temporal Action Detection
by: Liu, Shuming, et al.
Published: (2025)
by: Liu, Shuming, et al.
Published: (2025)
Skeleton-OOD: An End-to-End Skeleton-Based Model for Robust Out-of-Distribution Human Action Detection
by: Xu, Jing, et al.
Published: (2024)
by: Xu, Jing, et al.
Published: (2024)
LoSA: Long-Short-range Adapter for Scaling End-to-End Temporal Action Localization
by: Gupta, Akshita, et al.
Published: (2024)
by: Gupta, Akshita, et al.
Published: (2024)
YOLOv10: Real-Time End-to-End Object Detection
by: Wang, Ao, et al.
Published: (2024)
by: Wang, Ao, et al.
Published: (2024)
TE-TAD: Towards Full End-to-End Temporal Action Detection via Time-Aligned Coordinate Expression
by: Kim, Ho-Joong, et al.
Published: (2024)
by: Kim, Ho-Joong, et al.
Published: (2024)
End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling
by: Liang, Jianxin, et al.
Published: (2024)
by: Liang, Jianxin, et al.
Published: (2024)
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception
by: Li, Xiaoyu, et al.
Published: (2025)
by: Li, Xiaoyu, et al.
Published: (2025)
EF-VI: Enhancing End-Frame Injection for Video Inbetweening
by: Chen, Liuhan, et al.
Published: (2025)
by: Chen, Liuhan, et al.
Published: (2025)
Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving
by: Wang, Linbo, et al.
Published: (2026)
by: Wang, Linbo, et al.
Published: (2026)
An Effective End-to-End Solution for Multimodal Action Recognition
by: Wang, Songping, et al.
Published: (2025)
by: Wang, Songping, et al.
Published: (2025)
ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving
by: Zhang, Jinqing, et al.
Published: (2026)
by: Zhang, Jinqing, et al.
Published: (2026)
Temporally Consistent Dynamic Scene Graphs: An End-to-End Approach for Action Tracklet Generation
by: Ruschel, Raphael, et al.
Published: (2024)
by: Ruschel, Raphael, et al.
Published: (2024)
Multimodal Action Diffusion for Robust End-to-End Autonomous Driving
by: Rodríguez-Vidal, Jorge Daniel, et al.
Published: (2026)
by: Rodríguez-Vidal, Jorge Daniel, et al.
Published: (2026)
ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
by: Fu, Haoyu, et al.
Published: (2025)
by: Fu, Haoyu, et al.
Published: (2025)
End-to-End 4D Heart Mesh Recovery Across Full-Stack and Sparse Cardiac MRI
by: Chen, Yihong, et al.
Published: (2025)
by: Chen, Yihong, et al.
Published: (2025)
Towards End-to-End Explainable Facial Action Unit Recognition via Vision-Language Joint Learning
by: Ge, Xuri, et al.
Published: (2024)
by: Ge, Xuri, et al.
Published: (2024)
ColorMAE: Exploring data-independent masking strategies in Masked AutoEncoders
by: Hinojosa, Carlos, et al.
Published: (2024)
by: Hinojosa, Carlos, et al.
Published: (2024)
AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture-of-Transformers for End-to-End Autonomous Driving
by: Huang, Wenhui, et al.
Published: (2026)
by: Huang, Wenhui, et al.
Published: (2026)
Single Point, Full Mask: Velocity-Guided Level Set Evolution for End-to-End Amodal Segmentation
by: Li, Zhixuan, et al.
Published: (2025)
by: Li, Zhixuan, et al.
Published: (2025)
PAVE: An End-to-End Dataset for Production Autonomous Vehicle Evaluation
by: Li, Xiangyu, et al.
Published: (2025)
by: Li, Xiangyu, et al.
Published: (2025)
Manipulation Facing Threats: Evaluating Physical Vulnerabilities in End-to-End Vision Language Action Models
by: Cheng, Hao, et al.
Published: (2024)
by: Cheng, Hao, et al.
Published: (2024)
Efficient and Explainable End-to-End Autonomous Driving via Masked Vision-Language-Action Diffusion
by: Zhang, Jiaru, et al.
Published: (2026)
by: Zhang, Jiaru, et al.
Published: (2026)
GenAD: Generative End-to-End Autonomous Driving
by: Zheng, Wenzhao, et al.
Published: (2024)
by: Zheng, Wenzhao, et al.
Published: (2024)
End-to-End Spatial-Temporal Transformer for Real-time 4D HOI Reconstruction
by: Zhang, Haoyu, et al.
Published: (2026)
by: Zhang, Haoyu, et al.
Published: (2026)
UAV-DETR: Efficient End-to-End Object Detection for Unmanned Aerial Vehicle Imagery
by: Zhang, Huaxiang, et al.
Published: (2025)
by: Zhang, Huaxiang, et al.
Published: (2025)
End-to-End Action Segmentation Transformer
by: Wang, Tieqiao, et al.
Published: (2025)
by: Wang, Tieqiao, et al.
Published: (2025)
RQFormer: Rotated Query Transformer for End-to-End Oriented Object Detection
by: Zhao, Jiaqi, et al.
Published: (2023)
by: Zhao, Jiaqi, et al.
Published: (2023)
Action Images: End-to-End Policy Learning via Multiview Video Generation
by: Zhen, Haoyu, et al.
Published: (2026)
by: Zhen, Haoyu, et al.
Published: (2026)
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
by: Zhang, Mengchen, et al.
Published: (2025)
by: Zhang, Mengchen, et al.
Published: (2025)
From Human Intention to Action Prediction: Intention-Driven End-to-End Autonomous Driving
by: Zheng, Huan, et al.
Published: (2025)
by: Zheng, Huan, et al.
Published: (2025)
OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models
by: Zhong, Yufeng, et al.
Published: (2026)
by: Zhong, Yufeng, et al.
Published: (2026)
Polar R-CNN: End-to-End Lane Detection with Fewer Anchors
by: Wang, Shengqi, et al.
Published: (2024)
by: Wang, Shengqi, et al.
Published: (2024)
WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model
by: Zhang, Songyan, et al.
Published: (2024)
by: Zhang, Songyan, et al.
Published: (2024)
DEYO: DETR with YOLO for End-to-End Object Detection
by: Ouyang, Haodong
Published: (2024)
by: Ouyang, Haodong
Published: (2024)
End-to-End Facial Expression Detection in Long Videos
by: Fang, Yini, et al.
Published: (2025)
by: Fang, Yini, et al.
Published: (2025)
Similar Items
-
Harnessing Temporal Causality for Advanced Temporal Action Detection
by: Liu, Shuming, et al.
Published: (2024) -
TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos
by: Zhang, Chen-Lin, et al.
Published: (2025) -
Video Self-Stitching Graph Network for Temporal Action Localization
by: Zhao, Chen, et al.
Published: (2020) -
BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding
by: Liu, Shuming, et al.
Published: (2025) -
End-to-End Streaming Video Temporal Action Segmentation with Reinforce Learning
by: Zhang, Jinrong, et al.
Published: (2023)