A Semantic and Motion-Aware Spatiotemporal Transformer Network for Action Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Korban, Matthew, Youngs, Peter, Acton, Scott T. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
by: Kumar, Pulkit, et al.
Published: (2025)
by: Kumar, Pulkit, et al.
Published: (2025)
Capturing Rich Behavior Representations: A Dynamic Action Semantic-Aware Graph Transformer for Video Captioning
by: Liu, Caihua, et al.
Published: (2025)
by: Liu, Caihua, et al.
Published: (2025)
Uncertainty-Guided Appearance-Motion Association Network for Out-of-Distribution Action Detection
by: Fang, Xiang, et al.
Published: (2024)
by: Fang, Xiang, et al.
Published: (2024)
MoMaps: Semantics-Aware Scene Motion Generation with Motion Maps
by: Lei, Jiahui, et al.
Published: (2025)
by: Lei, Jiahui, et al.
Published: (2025)
Motion-Aware Transformer for Multi-Object Tracking
by: Yang, Xu, et al.
Published: (2025)
by: Yang, Xu, et al.
Published: (2025)
ActionPlan: Future-Aware Streaming Motion Synthesis via Frame-Level Action Planning
by: Nazarenus, Eric, et al.
Published: (2026)
by: Nazarenus, Eric, et al.
Published: (2026)
Multi-Stage Boundary-Aware Transformer Network for Action Segmentation in Untrimmed Surgical Videos
by: Shuvo, Rezowan, et al.
Published: (2025)
by: Shuvo, Rezowan, et al.
Published: (2025)
A Renaissance of Explicit Motion Information Mining from Transformers for Action Recognition
by: Zhuang, Peiqin, et al.
Published: (2025)
by: Zhuang, Peiqin, et al.
Published: (2025)
Pose-Aware Multi-Level Motion Parsing for Action Quality Assessment
by: Zhu, Shuaikang, et al.
Published: (2025)
by: Zhu, Shuaikang, et al.
Published: (2025)
Seeing Space and Motion: Enhancing Latent Actions with Geometric and Dynamic Awareness for Vision-Language-Action Models
by: Cai, Zhejia, et al.
Published: (2025)
by: Cai, Zhejia, et al.
Published: (2025)
FSGNet: A Frequency-Aware and Semantic Guidance Network for Infrared Small Target Detection
by: Zhang, Yingmei, et al.
Published: (2026)
by: Zhang, Yingmei, et al.
Published: (2026)
Lighting in Motion: Spatiotemporal HDR Lighting Estimation
by: Bolduc, Christophe, et al.
Published: (2025)
by: Bolduc, Christophe, et al.
Published: (2025)
Motion Matters: Motion-guided Modulation Network for Skeleton-based Micro-Action Recognition
by: Gu, Jihao, et al.
Published: (2025)
by: Gu, Jihao, et al.
Published: (2025)
Motion-Compensated Latent Semantic Canvases for Visual Situational Awareness on Edge
by: Lodin, Igor, et al.
Published: (2025)
by: Lodin, Igor, et al.
Published: (2025)
Future-Aware Interaction Network For Motion Forecasting
by: Li, Shijie, et al.
Published: (2025)
by: Li, Shijie, et al.
Published: (2025)
SemanticDialect: Semantic-Aware Mixed-Format Quantization for Video Diffusion Transformers
by: Jang, Wonsuk, et al.
Published: (2026)
by: Jang, Wonsuk, et al.
Published: (2026)
Unmasking Deepfakes: Masked Autoencoding Spatiotemporal Transformers for Enhanced Video Forgery Detection
by: Das, Sayantan, et al.
Published: (2023)
by: Das, Sayantan, et al.
Published: (2023)
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection
by: Bao, Wentao, et al.
Published: (2024)
by: Bao, Wentao, et al.
Published: (2024)
Bridging the Gap between Human Motion and Action Semantics via Kinematic Phrases
by: Liu, Xinpeng, et al.
Published: (2023)
by: Liu, Xinpeng, et al.
Published: (2023)
Semantics-Aware Human Motion Generation from Audio Instructions
by: Wang, Zi-An, et al.
Published: (2025)
by: Wang, Zi-An, et al.
Published: (2025)
Context and Geometry Aware Voxel Transformer for Semantic Scene Completion
by: Yu, Zhu, et al.
Published: (2024)
by: Yu, Zhu, et al.
Published: (2024)
Spatiotemporal-Untrammelled Mixture of Experts for Multi-Person Motion Prediction
by: Yin, Zheng, et al.
Published: (2025)
by: Yin, Zheng, et al.
Published: (2025)
Orientation-Aware Leg Movement Learning for Action-Driven Human Motion Prediction
by: Gu, Chunzhi, et al.
Published: (2023)
by: Gu, Chunzhi, et al.
Published: (2023)
Semantic-Aware Ship Detection with Vision-Language Integration
by: Li, Jiahao, et al.
Published: (2025)
by: Li, Jiahao, et al.
Published: (2025)
Video Anomaly Detection with Semantics-Aware Information Bottleneck
by: Li, Juntong, et al.
Published: (2025)
by: Li, Juntong, et al.
Published: (2025)
Long-term Pre-training for Temporal Action Detection with Transformers
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
Context-Enhanced Memory-Refined Transformer for Online Action Detection
by: Pang, Zhanzhong, et al.
Published: (2025)
by: Pang, Zhanzhong, et al.
Published: (2025)
Adapting Short-Term Transformers for Action Detection in Untrimmed Videos
by: Yang, Min, et al.
Published: (2023)
by: Yang, Min, et al.
Published: (2023)
Frequency Guidance Matters: Skeletal Action Recognition by Frequency-Aware Mixed Transformer
by: Wu, Wenhan, et al.
Published: (2024)
by: Wu, Wenhan, et al.
Published: (2024)
Hierarchical Multi-Stage Transformer Architecture for Context-Aware Temporal Action Localization
by: Ullah, Hayat, et al.
Published: (2025)
by: Ullah, Hayat, et al.
Published: (2025)
Boundary-Recovering Network for Temporal Action Detection
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
TAFormer: A Unified Target-Aware Transformer for Video and Motion Joint Prediction in Aerial Scenes
by: Xu, Liangyu, et al.
Published: (2024)
by: Xu, Liangyu, et al.
Published: (2024)
Motion Semantics Guided Normalizing Flow for Privacy-Preserving Video Anomaly Detection
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
Fine-Grained Spatiotemporal Motion Alignment for Contrastive Video Representation Learning
by: Zhu, Minghao, et al.
Published: (2023)
by: Zhu, Minghao, et al.
Published: (2023)
EgoAction: Egocentric Action Composition with Reliability-Aware Temporal Fusion for the EPIC-KITCHENS Action Detection Challenge at CVPR 2026
by: Fu, Zhiheng, et al.
Published: (2026)
by: Fu, Zhiheng, et al.
Published: (2026)
Context-Aware Interaction Network for RGB-T Semantic Segmentation
by: Lv, Ying, et al.
Published: (2024)
by: Lv, Ying, et al.
Published: (2024)
Occlusion-Aware 3D Motion Interpretation for Abnormal Behavior Detection
by: Li, Su, et al.
Published: (2024)
by: Li, Su, et al.
Published: (2024)
Unsupervised 4D Cardiac Motion Tracking with Spatiotemporal Optical Flow Networks
by: Teng, Long, et al.
Published: (2024)
by: Teng, Long, et al.
Published: (2024)
Flow Snapshot Neurons in Action: Deep Neural Networks Generalize to Biological Motion Perception
by: Han, Shuangpeng, et al.
Published: (2024)
by: Han, Shuangpeng, et al.
Published: (2024)
Motion-Adapter: A Diffusion Model Adapter for Text-to-Motion Generation of Compound Actions
by: Jiang, Yue, et al.
Published: (2026)
by: Jiang, Yue, et al.
Published: (2026)
Similar Items
-
Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
by: Kumar, Pulkit, et al.
Published: (2025) -
Capturing Rich Behavior Representations: A Dynamic Action Semantic-Aware Graph Transformer for Video Captioning
by: Liu, Caihua, et al.
Published: (2025) -
Uncertainty-Guided Appearance-Motion Association Network for Out-of-Distribution Action Detection
by: Fang, Xiang, et al.
Published: (2024) -
MoMaps: Semantics-Aware Scene Motion Generation with Motion Maps
by: Lei, Jiahui, et al.
Published: (2025) -
Motion-Aware Transformer for Multi-Object Tracking
by: Yang, Xu, et al.
Published: (2025)