Context-Enhanced Memory-Refined Transformer for Online Action Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Pang, Zhanzhong, Sener, Fadime, Yao, Angela |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cost-Sensitive Learning for Long-Tailed Temporal Action Segmentation
by: Pang, Zhanzhong, et al.
Published: (2025)
by: Pang, Zhanzhong, et al.
Published: (2025)
On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action Understanding
by: Pang, Zhanzhong, et al.
Published: (2026)
by: Pang, Zhanzhong, et al.
Published: (2026)
Long-Tail Temporal Action Segmentation with Group-wise Temporal Logit Adjustment
by: Pang, Zhanzhong, et al.
Published: (2024)
by: Pang, Zhanzhong, et al.
Published: (2024)
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
by: Pang, Zhanzhong, et al.
Published: (2026)
by: Pang, Zhanzhong, et al.
Published: (2026)
Don't Pause! Every prediction matters in a streaming video
by: Chatterjee, Dibyadip, et al.
Published: (2026)
by: Chatterjee, Dibyadip, et al.
Published: (2026)
On the Utility of 3D Hand Poses for Action Recognition
by: Shamil, Md Salman, et al.
Published: (2024)
by: Shamil, Md Salman, et al.
Published: (2024)
X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization
by: Kukleva, Anna, et al.
Published: (2024)
by: Kukleva, Anna, et al.
Published: (2024)
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
by: Chatterjee, Dibyadip, et al.
Published: (2025)
by: Chatterjee, Dibyadip, et al.
Published: (2025)
OnlineTAS: An Online Baseline for Temporal Action Segmentation
by: Zhong, Qing, et al.
Published: (2024)
by: Zhong, Qing, et al.
Published: (2024)
Online Temporal Action Localization with Memory-Augmented Transformer
by: Song, Youngkil, et al.
Published: (2024)
by: Song, Youngkil, et al.
Published: (2024)
PALM: A Dataset and Baseline for Learning Multi-subject Hand Prior
by: Fan, Zicong, et al.
Published: (2025)
by: Fan, Zicong, et al.
Published: (2025)
Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering Refinement
by: Hou, Xiuquan, et al.
Published: (2024)
by: Hou, Xiuquan, et al.
Published: (2024)
MALT: Multi-scale Action Learning Transformer for Online Action Detection
by: Yang, Zhipeng, et al.
Published: (2024)
by: Yang, Zhipeng, et al.
Published: (2024)
SneakPeek: Future-Guided Instructional Streaming Video Generation
by: Hong, Cheeun, et al.
Published: (2025)
by: Hong, Cheeun, et al.
Published: (2025)
Text-driven Online Action Detection
by: Benavent-Lledo, Manuel, et al.
Published: (2025)
by: Benavent-Lledo, Manuel, et al.
Published: (2025)
DetRefiner: Model-Agnostic Detection Refinement with Feature Fusion Transformer
by: Okazaki, Soichiro, et al.
Published: (2026)
by: Okazaki, Soichiro, et al.
Published: (2026)
FDDet: Frequency-Decoupling for Boundary Refinement in Temporal Action Detection
by: Zhu, Xinnan, et al.
Published: (2025)
by: Zhu, Xinnan, et al.
Published: (2025)
Introducing Gating and Context into Temporal Action Detection
by: Reka, Aglind, et al.
Published: (2024)
by: Reka, Aglind, et al.
Published: (2024)
Unveiling Context-Related Anomalies: Knowledge Graph Empowered Decoupling of Scene and Action for Human-Related Video Anomaly Detection
by: Chen, Chenglizhao, et al.
Published: (2024)
by: Chen, Chenglizhao, et al.
Published: (2024)
Track-On: Transformer-based Online Point Tracking with Memory
by: Aydemir, Görkay, et al.
Published: (2025)
by: Aydemir, Görkay, et al.
Published: (2025)
OMR: Occlusion-Aware Memory-Based Refinement for Video Lane Detection
by: Jin, Dongkwon, et al.
Published: (2024)
by: Jin, Dongkwon, et al.
Published: (2024)
KITRO: Refining Human Mesh by 2D Clues and Kinematic-tree Rotation
by: Yang, Fengyuan, et al.
Published: (2024)
by: Yang, Fengyuan, et al.
Published: (2024)
Information Elevation Network for Fast Online Action Detection
by: Min, Sunah, et al.
Published: (2021)
by: Min, Sunah, et al.
Published: (2021)
Coherent Temporal Synthesis for Incremental Action Segmentation
by: Ding, Guodong, et al.
Published: (2024)
by: Ding, Guodong, et al.
Published: (2024)
Retrospective Memory for Camouflaged Object Detection
by: Zhang, Chenxi, et al.
Published: (2025)
by: Zhang, Chenxi, et al.
Published: (2025)
Track-On2: Enhancing Online Point Tracking with Memory
by: Aydemir, Görkay, et al.
Published: (2025)
by: Aydemir, Görkay, et al.
Published: (2025)
HAT: History-Augmented Anchor Transformer for Online Temporal Action Localization
by: Reza, Sakib, et al.
Published: (2024)
by: Reza, Sakib, et al.
Published: (2024)
Condensing Action Segmentation Datasets via Generative Network Inversion
by: Ding, Guodong, et al.
Published: (2025)
by: Ding, Guodong, et al.
Published: (2025)
Online Action Representation using Change Detection and Symbolic Programming
by: Nair, Vishnu S, et al.
Published: (2024)
by: Nair, Vishnu S, et al.
Published: (2024)
Towards Online Real-Time Memory-based Video Inpainting Transformers
by: Thiry, Guillaume, et al.
Published: (2024)
by: Thiry, Guillaume, et al.
Published: (2024)
Modeling Multi-Granularity Context Information Flow for Pavement Crack Detection
by: Pang, Junbiao, et al.
Published: (2024)
by: Pang, Junbiao, et al.
Published: (2024)
Uncertainty-Masked Bernoulli Diffusion for Camouflaged Object Detection Refinement
by: Shen, Yuqi, et al.
Published: (2025)
by: Shen, Yuqi, et al.
Published: (2025)
Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer
by: Xing, Bohao, et al.
Published: (2026)
by: Xing, Bohao, et al.
Published: (2026)
Hierarchical Multi-Stage Transformer Architecture for Context-Aware Temporal Action Localization
by: Ullah, Hayat, et al.
Published: (2025)
by: Ullah, Hayat, et al.
Published: (2025)
Crane: Context-Guided Prompt Learning and Attention Refinement for Zero-Shot Anomaly Detection
by: Salehi, Alireza, et al.
Published: (2025)
by: Salehi, Alireza, et al.
Published: (2025)
Enhancing Video Transformers for Action Understanding with VLM-aided Training
by: Lu, Hui, et al.
Published: (2024)
by: Lu, Hui, et al.
Published: (2024)
Uncertainty Guided Refinement for Fine-Grained Salient Object Detection
by: Yuan, Yao, et al.
Published: (2025)
by: Yuan, Yao, et al.
Published: (2025)
Adapting Short-Term Transformers for Action Detection in Untrimmed Videos
by: Yang, Min, et al.
Published: (2023)
by: Yang, Min, et al.
Published: (2023)
Long-term Pre-training for Temporal Action Detection with Transformers
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
Few-Shot Object Detection with Sparse Context Transformers
by: Mei, Jie, et al.
Published: (2024)
by: Mei, Jie, et al.
Published: (2024)
Similar Items
-
Cost-Sensitive Learning for Long-Tailed Temporal Action Segmentation
by: Pang, Zhanzhong, et al.
Published: (2025) -
On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action Understanding
by: Pang, Zhanzhong, et al.
Published: (2026) -
Long-Tail Temporal Action Segmentation with Group-wise Temporal Logit Adjustment
by: Pang, Zhanzhong, et al.
Published: (2024) -
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
by: Pang, Zhanzhong, et al.
Published: (2026) -
Don't Pause! Every prediction matters in a streaming video
by: Chatterjee, Dibyadip, et al.
Published: (2026)