Transformer-based Fusion of 2D-pose and Spatio-temporal Embeddings for Distracted Driver Action Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Akdag, Erkut, Zhu, Zeqi, Bondarev, Egor, De With, Peter H. N. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Density-Guided Label Smoothing for Temporal Localization of Driving Actions
by: Alkanat, Tunc, et al.
Published: (2024)
by: Alkanat, Tunc, et al.
Published: (2024)
TeG: Temporal-Granularity Method for Anomaly Detection with Attention in Smart City Surveillance
by: Akdag, Erkut, et al.
Published: (2024)
by: Akdag, Erkut, et al.
Published: (2024)
Boxes2Pixels: Learning Defect Segmentation from Noisy SAM Masks
by: Lendering, Camile, et al.
Published: (2026)
by: Lendering, Camile, et al.
Published: (2026)
MTFL: Multi-Timescale Feature Learning for Weakly-Supervised Anomaly Detection in Surveillance Videos
by: Zhang, Yiling, et al.
Published: (2024)
by: Zhang, Yiling, et al.
Published: (2024)
SubspaceAD: Training-Free Few-Shot Anomaly Detection via Subspace Modeling
by: Lendering, Camile, et al.
Published: (2026)
by: Lendering, Camile, et al.
Published: (2026)
Detection of Object Throwing Behavior in Surveillance Videos
by: Kersten, Ivo P. C., et al.
Published: (2024)
by: Kersten, Ivo P. C., et al.
Published: (2024)
LiveStre4m: Feed-Forward Live Streaming of Novel Views from Unposed Multi-View Video
by: Quesado, Pedro, et al.
Published: (2026)
by: Quesado, Pedro, et al.
Published: (2026)
F3G-Avatar : Face Focused Full-body Gaussian Avatar
by: Menu, Willem, et al.
Published: (2026)
by: Menu, Willem, et al.
Published: (2026)
R3PM-Net: Real-time, Robust, Real-world Point Matching Network
by: Kashefbahrami, Yasaman, et al.
Published: (2026)
by: Kashefbahrami, Yasaman, et al.
Published: (2026)
MultiFuser: Multimodal Fusion Transformer for Enhanced Driver Action Recognition
by: Wang, Ruoyu, et al.
Published: (2024)
by: Wang, Ruoyu, et al.
Published: (2024)
Unmasking Performance Gaps: A Comparative Study of Human Anonymization and Its Effects on Video Anomaly Detection
by: Abdulaziz, Sara, et al.
Published: (2025)
by: Abdulaziz, Sara, et al.
Published: (2025)
MIFI: MultI-camera Feature Integration for Roust 3D Distracted Driver Activity Recognition
by: Kuang, Jian, et al.
Published: (2024)
by: Kuang, Jian, et al.
Published: (2024)
Spatio-temporal Decoupled Knowledge Compensator for Few-Shot Action Recognition
by: Qu, Hongyu, et al.
Published: (2026)
by: Qu, Hongyu, et al.
Published: (2026)
Spatio-temporal Transformers for Action Unit Classification with Event Cameras
by: Cultrera, Luca, et al.
Published: (2024)
by: Cultrera, Luca, et al.
Published: (2024)
MVAFormer: RGB-based Multi-View Spatio-Temporal Action Recognition with Transformer
by: Yamane, Taiga, et al.
Published: (2025)
by: Yamane, Taiga, et al.
Published: (2025)
ViT-DD: Multi-Task Vision Transformer for Semi-Supervised Driver Distraction Detection
by: Ma, Yunsheng, et al.
Published: (2022)
by: Ma, Yunsheng, et al.
Published: (2022)
DSDFormer: An Innovative Transformer-Mamba Framework for Robust High-Precision Driver Distraction Identification
by: Chen, Junzhou, et al.
Published: (2024)
by: Chen, Junzhou, et al.
Published: (2024)
Spiking-DD: Neuromorphic Event Camera based Driver Distraction Detection with Spiking Neural Network
by: Shariff, Waseem, et al.
Published: (2024)
by: Shariff, Waseem, et al.
Published: (2024)
Evaluation of Human Visual Privacy Protection: A Three-Dimensional Framework and Benchmark Dataset
by: Abdulaziz, Sara, et al.
Published: (2025)
by: Abdulaziz, Sara, et al.
Published: (2025)
Otter: Mitigating Background Distractions of Wide-Angle Few-Shot Action Recognition with Enhanced RWKV
by: Huang, Wenbo, et al.
Published: (2025)
by: Huang, Wenbo, et al.
Published: (2025)
Latent Uncertainty Representations for Video-based Driver Action and Intention Recognition
by: Vellenga, Koen, et al.
Published: (2025)
by: Vellenga, Koen, et al.
Published: (2025)
UniSTFormer: Unified Spatio-Temporal Lightweight Transformer for Efficient Skeleton-Based Action Recognition
by: Wu, Wenhan, et al.
Published: (2025)
by: Wu, Wenhan, et al.
Published: (2025)
3DPyranet Features Fusion for Spatio-temporal Feature Learning
by: Ullah, Ihsan, et al.
Published: (2025)
by: Ullah, Ihsan, et al.
Published: (2025)
Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers
by: Claessens, Cris, et al.
Published: (2025)
by: Claessens, Cris, et al.
Published: (2025)
Modelling Spatio-Temporal Interactions For Compositional Action Recognition
by: Rajendiran, Ramanathan, et al.
Published: (2023)
by: Rajendiran, Ramanathan, et al.
Published: (2023)
Co-Fusion4D: Spatio-temporal Collaborative Fusion for Robust 3D Object Detection
by: Li, Wenxuan, et al.
Published: (2026)
by: Li, Wenxuan, et al.
Published: (2026)
MultiTSF: Transformer-based Sensor Fusion for Human-Centric Multi-view and Multi-modal Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2025)
by: Nguyen, Trung Thanh, et al.
Published: (2025)
Vision-Language Models can Identify Distracted Driver Behavior from Naturalistic Videos
by: Hasan, Md Zahid, et al.
Published: (2023)
by: Hasan, Md Zahid, et al.
Published: (2023)
Multi-modal Video Representation Alignment for Robust Self-supervised Driver Distraction Detection
by: Lerch, David J., et al.
Published: (2026)
by: Lerch, David J., et al.
Published: (2026)
Blur-aware Spatio-temporal Sparse Transformer for Video Deblurring
by: Zhang, Huicong, et al.
Published: (2024)
by: Zhang, Huicong, et al.
Published: (2024)
Spatio-Temporal Similarity Volume Aggregation for Open-Vocabulary Action Recognition
by: So, Yerim, et al.
Published: (2026)
by: So, Yerim, et al.
Published: (2026)
Enhancing Road Safety: Real-Time Detection of Driver Distraction through Convolutional Neural Networks
by: Sheikh, Amaan Aijaz, et al.
Published: (2024)
by: Sheikh, Amaan Aijaz, et al.
Published: (2024)
EyeCue: Driver Cognitive Distraction Detection via Gaze-Empowered Egocentric Video Understanding
by: Zhang, Lang, et al.
Published: (2026)
by: Zhang, Lang, et al.
Published: (2026)
D$^2$ST-Adapter: Disentangled-and-Deformable Spatio-Temporal Adapter for Few-shot Action Recognition
by: Pei, Wenjie, et al.
Published: (2023)
by: Pei, Wenjie, et al.
Published: (2023)
Dual-view Spatio-Temporal Feature Fusion with CNN-Transformer Hybrid Network for Chinese Isolated Sign Language Recognition
by: Jing, Siyuan, et al.
Published: (2025)
by: Jing, Siyuan, et al.
Published: (2025)
LORTSAR: Low-Rank Transformer for Skeleton-based Action Recognition
by: Oraki, Soroush, et al.
Published: (2024)
by: Oraki, Soroush, et al.
Published: (2024)
FineParser: A Fine-grained Spatio-temporal Action Parser for Human-centric Action Quality Assessment
by: Xu, Jinglin, et al.
Published: (2024)
by: Xu, Jinglin, et al.
Published: (2024)
Game State and Spatio-temporal Action Detection in Soccer using Graph Neural Networks and 3D Convolutional Networks
by: Ochin, Jeremie, et al.
Published: (2025)
by: Ochin, Jeremie, et al.
Published: (2025)
VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
Spatio-Temporal Joint Density Driven Learning for Skeleton-Based Action Recognition
by: Gunasekara, Shanaka Ramesh, et al.
Published: (2025)
by: Gunasekara, Shanaka Ramesh, et al.
Published: (2025)
Similar Items
-
Density-Guided Label Smoothing for Temporal Localization of Driving Actions
by: Alkanat, Tunc, et al.
Published: (2024) -
TeG: Temporal-Granularity Method for Anomaly Detection with Attention in Smart City Surveillance
by: Akdag, Erkut, et al.
Published: (2024) -
Boxes2Pixels: Learning Defect Segmentation from Noisy SAM Masks
by: Lendering, Camile, et al.
Published: (2026) -
MTFL: Multi-Timescale Feature Learning for Weakly-Supervised Anomaly Detection in Surveillance Videos
by: Zhang, Yiling, et al.
Published: (2024) -
SubspaceAD: Training-Free Few-Shot Anomaly Detection via Subspace Modeling
by: Lendering, Camile, et al.
Published: (2026)