TE-TAD: Towards Full End-to-End Temporal Action Detection via Time-Aligned Coordinate Expression
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Ho-Joong, Hong, Jung-Ho, Kong, Heejo, Lee, Seong-Whan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection Transformer
di: Kim, Ho-Joong, et al.
Pubblicazione: (2025)
di: Kim, Ho-Joong, et al.
Pubblicazione: (2025)
Text-guided Weakly Supervised Framework for Dynamic Facial Expression Recognition
di: Jung, Gunho, et al.
Pubblicazione: (2025)
di: Jung, Gunho, et al.
Pubblicazione: (2025)
Comprehensive Information Bottleneck for Unveiling Universal Attribution to Interpret Vision Transformers
di: Hong, Jung-Ho, et al.
Pubblicazione: (2025)
di: Hong, Jung-Ho, et al.
Pubblicazione: (2025)
Diversify and Conquer: Open-set Disagreement for Robust Semi-supervised Learning with Outliers
di: Kong, Heejo, et al.
Pubblicazione: (2025)
di: Kong, Heejo, et al.
Pubblicazione: (2025)
ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval
di: Kim, Ji-Hyeon, et al.
Pubblicazione: (2026)
di: Kim, Ji-Hyeon, et al.
Pubblicazione: (2026)
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
di: Ju, Yeong-Joon, et al.
Pubblicazione: (2024)
di: Ju, Yeong-Joon, et al.
Pubblicazione: (2024)
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
di: Oh, Ju-Young, et al.
Pubblicazione: (2025)
di: Oh, Ju-Young, et al.
Pubblicazione: (2025)
Towards Better Visualizing the Decision Basis of Networks via Unfold and Conquer Attribution Guidance
di: Hong, Jung-Ho, et al.
Pubblicazione: (2023)
di: Hong, Jung-Ho, et al.
Pubblicazione: (2023)
Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes
di: Kim, Yehna, et al.
Pubblicazione: (2025)
di: Kim, Yehna, et al.
Pubblicazione: (2025)
LiquidTAD: Efficient Temporal Action Detection via Parallel Liquid-Inspired Temporal Relaxation
di: Sun, Zepeng, et al.
Pubblicazione: (2026)
di: Sun, Zepeng, et al.
Pubblicazione: (2026)
CREPE: Coordinate-Aware End-to-End Document Parser
di: Okamoto, Yamato, et al.
Pubblicazione: (2024)
di: Okamoto, Yamato, et al.
Pubblicazione: (2024)
End-to-End Temporal Action Detection with 1B Parameters Across 1000 Frames
di: Liu, Shuming, et al.
Pubblicazione: (2023)
di: Liu, Shuming, et al.
Pubblicazione: (2023)
Towards End-to-End Semi-Supervised Table Detection with Semantic Aligned Matching Transformer
di: Shehzadi, Tahira, et al.
Pubblicazione: (2024)
di: Shehzadi, Tahira, et al.
Pubblicazione: (2024)
Prediction-Feedback DETR for Temporal Action Detection
di: Kim, Jihwan, et al.
Pubblicazione: (2024)
di: Kim, Jihwan, et al.
Pubblicazione: (2024)
AM-SORT: Adaptable Motion Predictor with Historical Trajectory Embedding for Multi-Object Tracking
di: Kim, Vitaliy, et al.
Pubblicazione: (2024)
di: Kim, Vitaliy, et al.
Pubblicazione: (2024)
End-to-End Facial Expression Detection in Long Videos
di: Fang, Yini, et al.
Pubblicazione: (2025)
di: Fang, Yini, et al.
Pubblicazione: (2025)
OpenTAD: A Unified Framework and Comprehensive Study of Temporal Action Detection
di: Liu, Shuming, et al.
Pubblicazione: (2025)
di: Liu, Shuming, et al.
Pubblicazione: (2025)
End-to-End Streaming Video Temporal Action Segmentation with Reinforce Learning
di: Zhang, Jinrong, et al.
Pubblicazione: (2023)
di: Zhang, Jinrong, et al.
Pubblicazione: (2023)
VLA-R: Vision-Language Action Retrieval toward Open-World End-to-End Autonomous Driving
di: Seong, Hyunki, et al.
Pubblicazione: (2025)
di: Seong, Hyunki, et al.
Pubblicazione: (2025)
Towards End-to-End Explainable Facial Action Unit Recognition via Vision-Language Joint Learning
di: Ge, Xuri, et al.
Pubblicazione: (2024)
di: Ge, Xuri, et al.
Pubblicazione: (2024)
Safety-Aligned 3D Object Detection: Single-Vehicle, Cooperative, and End-to-End Perspectives
di: Liao, Brian Hsuan-Cheng, et al.
Pubblicazione: (2026)
di: Liao, Brian Hsuan-Cheng, et al.
Pubblicazione: (2026)
LoSA: Long-Short-range Adapter for Scaling End-to-End Temporal Action Localization
di: Gupta, Akshita, et al.
Pubblicazione: (2024)
di: Gupta, Akshita, et al.
Pubblicazione: (2024)
Align-DETR: Enhancing End-to-end Object Detection with Aligned Loss
di: Cai, Zhi, et al.
Pubblicazione: (2023)
di: Cai, Zhi, et al.
Pubblicazione: (2023)
Appearance Debiased Gaze Estimation via Stochastic Subject-Wise Adversarial Learning
di: Kim, Suneung, et al.
Pubblicazione: (2024)
di: Kim, Suneung, et al.
Pubblicazione: (2024)
Multi-Context Temporal Consistent Modeling for Referring Video Object Segmentation
di: Choi, Sun-Hyuk, et al.
Pubblicazione: (2025)
di: Choi, Sun-Hyuk, et al.
Pubblicazione: (2025)
MambaTAD: When State-Space Models Meet Long-Range Temporal Action Detection
di: Lu, Hui, et al.
Pubblicazione: (2025)
di: Lu, Hui, et al.
Pubblicazione: (2025)
AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving
di: Wu, Yanhao, et al.
Pubblicazione: (2026)
di: Wu, Yanhao, et al.
Pubblicazione: (2026)
Temporally Consistent Dynamic Scene Graphs: An End-to-End Approach for Action Tracklet Generation
di: Ruschel, Raphael, et al.
Pubblicazione: (2024)
di: Ruschel, Raphael, et al.
Pubblicazione: (2024)
YOLOv10: Real-Time End-to-End Object Detection
di: Wang, Ao, et al.
Pubblicazione: (2024)
di: Wang, Ao, et al.
Pubblicazione: (2024)
An Effective End-to-End Solution for Multimodal Action Recognition
di: Wang, Songping, et al.
Pubblicazione: (2025)
di: Wang, Songping, et al.
Pubblicazione: (2025)
TIFu: Tri-directional Implicit Function for High-Fidelity 3D Character Reconstruction
di: Lim, Byoungsung, et al.
Pubblicazione: (2024)
di: Lim, Byoungsung, et al.
Pubblicazione: (2024)
End-to-End Action Segmentation Transformer
di: Wang, Tieqiao, et al.
Pubblicazione: (2025)
di: Wang, Tieqiao, et al.
Pubblicazione: (2025)
Action Images: End-to-End Policy Learning via Multiview Video Generation
di: Zhen, Haoyu, et al.
Pubblicazione: (2026)
di: Zhen, Haoyu, et al.
Pubblicazione: (2026)
Referring Expression Instance Retrieval and A Strong End-to-End Baseline
di: Hao, Xiangzhao, et al.
Pubblicazione: (2025)
di: Hao, Xiangzhao, et al.
Pubblicazione: (2025)
Multimodal Action Diffusion for Robust End-to-End Autonomous Driving
di: Rodríguez-Vidal, Jorge Daniel, et al.
Pubblicazione: (2026)
di: Rodríguez-Vidal, Jorge Daniel, et al.
Pubblicazione: (2026)
RaDL: Relation-aware Disentangled Learning for Multi-Instance Text-to-Image Generation
di: Park, Geon, et al.
Pubblicazione: (2025)
di: Park, Geon, et al.
Pubblicazione: (2025)
A Differentiable Wave Optics Model for End-to-End Computational Imaging System Optimization
di: Ho, Chi-Jui, et al.
Pubblicazione: (2024)
di: Ho, Chi-Jui, et al.
Pubblicazione: (2024)
Skeleton-OOD: An End-to-End Skeleton-Based Model for Robust Out-of-Distribution Human Action Detection
di: Xu, Jing, et al.
Pubblicazione: (2024)
di: Xu, Jing, et al.
Pubblicazione: (2024)
End-to-End Full-Page Optical Music Recognition for Pianoform Sheet Music
di: Ríos-Vila, Antonio, et al.
Pubblicazione: (2024)
di: Ríos-Vila, Antonio, et al.
Pubblicazione: (2024)
Efficient and Explainable End-to-End Autonomous Driving via Masked Vision-Language-Action Diffusion
di: Zhang, Jiaru, et al.
Pubblicazione: (2026)
di: Zhang, Jiaru, et al.
Pubblicazione: (2026)
Documenti analoghi
-
DiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection Transformer
di: Kim, Ho-Joong, et al.
Pubblicazione: (2025) -
Text-guided Weakly Supervised Framework for Dynamic Facial Expression Recognition
di: Jung, Gunho, et al.
Pubblicazione: (2025) -
Comprehensive Information Bottleneck for Unveiling Universal Attribution to Interpret Vision Transformers
di: Hong, Jung-Ho, et al.
Pubblicazione: (2025) -
Diversify and Conquer: Open-set Disagreement for Robust Semi-supervised Learning with Outliers
di: Kong, Heejo, et al.
Pubblicazione: (2025) -
ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval
di: Kim, Ji-Hyeon, et al.
Pubblicazione: (2026)