Online Temporal Action Localization with Memory-Augmented Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Youngkil, Kim, Dongkeun, Cho, Minsu, Kwak, Suha |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards More Practical Group Activity Detection: A New Benchmark and Model
by: Kim, Dongkeun, et al.
Published: (2023)
by: Kim, Dongkeun, et al.
Published: (2023)
Part-Aware Bottom-Up Group Reasoning for Fine-Grained Social Interaction Detection
by: Kim, Dongkeun, et al.
Published: (2025)
by: Kim, Dongkeun, et al.
Published: (2025)
Improving Sound Source Localization with Joint Slot Attention on Image and Audio
by: Kim, Inho, et al.
Published: (2025)
by: Kim, Inho, et al.
Published: (2025)
Tempered Self-Similarity Alignment for Physically Plausible Video Generation
by: Kim, Manjin, et al.
Published: (2026)
by: Kim, Manjin, et al.
Published: (2026)
ActFusion: a Unified Diffusion Model for Action Segmentation and Anticipation
by: Gong, Dayoung, et al.
Published: (2024)
by: Gong, Dayoung, et al.
Published: (2024)
Classification Matters: Improving Video Action Detection with Class-Specific Attention
by: Lee, Jinsung, et al.
Published: (2024)
by: Lee, Jinsung, et al.
Published: (2024)
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
by: Kim, Dongwon, et al.
Published: (2026)
by: Kim, Dongwon, et al.
Published: (2026)
HAT: History-Augmented Anchor Transformer for Online Temporal Action Localization
by: Reza, Sakib, et al.
Published: (2024)
by: Reza, Sakib, et al.
Published: (2024)
Learning SO(3)-Invariant Semantic Correspondence via Local Shape Transform
by: Park, Chunghyun, et al.
Published: (2024)
by: Park, Chunghyun, et al.
Published: (2024)
Locality-Aware Zero-Shot Human-Object Interaction Detection
by: Kim, Sanghyun, et al.
Published: (2025)
by: Kim, Sanghyun, et al.
Published: (2025)
Bootstrapping Top-down Information for Self-modulating Slot Attention
by: Kim, Dongwon, et al.
Published: (2024)
by: Kim, Dongwon, et al.
Published: (2024)
FreeAction: Training-Free Techniques for Enhanced Fidelity of Trajectory-to-Video Generation
by: Kim, Seungwook, et al.
Published: (2025)
by: Kim, Seungwook, et al.
Published: (2025)
FREST: Feature RESToration for Semantic Segmentation under Multiple Adverse Conditions
by: Lee, Sohyun, et al.
Published: (2024)
by: Lee, Sohyun, et al.
Published: (2024)
Active Label Correction for Semantic Segmentation with Foundation Models
by: Kim, Hoyoung, et al.
Published: (2024)
by: Kim, Hoyoung, et al.
Published: (2024)
PLOT: Text-based Person Search with Part Slot Attention for Corresponding Part Discovery
by: Park, Jicheol, et al.
Published: (2024)
by: Park, Jicheol, et al.
Published: (2024)
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
by: Jeong, Boseung, et al.
Published: (2025)
by: Jeong, Boseung, et al.
Published: (2025)
Similarity-Aware Selective State-Space Modeling for Semantic Correspondence
by: Kim, Seungwook, et al.
Published: (2025)
by: Kim, Seungwook, et al.
Published: (2025)
Extreme Point Supervised Instance Segmentation
by: Lee, Hyeonjun, et al.
Published: (2024)
by: Lee, Hyeonjun, et al.
Published: (2024)
RePL: Pseudo-label Refinement for Semi-supervised LiDAR Semantic Segmentation
by: Kwon, Donghyeon, et al.
Published: (2026)
by: Kwon, Donghyeon, et al.
Published: (2026)
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers
by: Bae, Jongseong, et al.
Published: (2024)
by: Bae, Jongseong, et al.
Published: (2024)
Learning Unified Distance Metric Across Diverse Data Distributions with Parameter-Efficient Transfer Learning
by: Kim, Sungyeon, et al.
Published: (2023)
by: Kim, Sungyeon, et al.
Published: (2023)
Learning Correlation Structures for Vision Transformers
by: Kim, Manjin, et al.
Published: (2024)
by: Kim, Manjin, et al.
Published: (2024)
Extending CLIP's Image-Text Alignment to Referring Image Segmentation
by: Kim, Seoyeon, et al.
Published: (2023)
by: Kim, Seoyeon, et al.
Published: (2023)
GroupCoOp: Group-robust Fine-tuning via Group Prompt Learning
by: Kim, Nayeong, et al.
Published: (2025)
by: Kim, Nayeong, et al.
Published: (2025)
Efficient and Versatile Robust Fine-Tuning of Zero-shot Models
by: Kim, Sungyeon, et al.
Published: (2024)
by: Kim, Sungyeon, et al.
Published: (2024)
OZ-TAL: Online Zero-Shot Temporal Action Localization
by: Han, Chaolei, et al.
Published: (2026)
by: Han, Chaolei, et al.
Published: (2026)
Context-Enhanced Memory-Refined Transformer for Online Action Detection
by: Pang, Zhanzhong, et al.
Published: (2025)
by: Pang, Zhanzhong, et al.
Published: (2025)
Improving Text-based Person Search via Part-level Cross-modal Correspondence
by: Park, Jicheol, et al.
Published: (2024)
by: Park, Jicheol, et al.
Published: (2024)
Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards
by: Kim, Seungwook, et al.
Published: (2026)
by: Kim, Seungwook, et al.
Published: (2026)
GaRA-SAM: Robustifying Segment Anything Model with Gated-Rank Adaptation
by: Lee, Sohyun, et al.
Published: (2025)
by: Lee, Sohyun, et al.
Published: (2025)
Harnessing the Power of Training-Free Techniques in Text-to-2D Generation for Text-to-3D Generation via Score Distillation Sampling
by: Lee, Junhong, et al.
Published: (2025)
by: Lee, Junhong, et al.
Published: (2025)
Memory-Modular Classification: Learning to Generalize with Memory Replacement
by: Kang, Dahyun, et al.
Published: (2025)
by: Kang, Dahyun, et al.
Published: (2025)
In Defense of Lazy Visual Grounding for Open-Vocabulary Semantic Segmentation
by: Kang, Dahyun, et al.
Published: (2024)
by: Kang, Dahyun, et al.
Published: (2024)
Leveraging 3D Geometric Priors in 2D Rotation Symmetry Detection
by: Seo, Ahyun, et al.
Published: (2025)
by: Seo, Ahyun, et al.
Published: (2025)
Structured State-Space Regularization for Generation-Friendly Image Tokenization
by: Lee, Jinsung, et al.
Published: (2026)
by: Lee, Jinsung, et al.
Published: (2026)
SkateFormer: Skeletal-Temporal Transformer for Human Action Recognition
by: Do, Jeonghyeok, et al.
Published: (2024)
by: Do, Jeonghyeok, et al.
Published: (2024)
Probabilistic Vision-Language Representation for Weakly Supervised Temporal Action Localization
by: Lim, Geuntaek, et al.
Published: (2024)
by: Lim, Geuntaek, et al.
Published: (2024)
OnlineTAS: An Online Baseline for Temporal Action Segmentation
by: Zhong, Qing, et al.
Published: (2024)
by: Zhong, Qing, et al.
Published: (2024)
Prediction-Feedback DETR for Temporal Action Detection
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
Hierarchically Structured Neural Bones for Reconstructing Animatable Objects from Casual Videos
by: Jeon, Subin, et al.
Published: (2024)
by: Jeon, Subin, et al.
Published: (2024)
Similar Items
-
Towards More Practical Group Activity Detection: A New Benchmark and Model
by: Kim, Dongkeun, et al.
Published: (2023) -
Part-Aware Bottom-Up Group Reasoning for Fine-Grained Social Interaction Detection
by: Kim, Dongkeun, et al.
Published: (2025) -
Improving Sound Source Localization with Joint Slot Attention on Image and Audio
by: Kim, Inho, et al.
Published: (2025) -
Tempered Self-Similarity Alignment for Physically Plausible Video Generation
by: Kim, Manjin, et al.
Published: (2026) -
ActFusion: a Unified Diffusion Model for Action Segmentation and Anticipation
by: Gong, Dayoung, et al.
Published: (2024)