Multi-modal Video Representation Alignment for Robust Self-supervised Driver Distraction Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Lerch, David J., Majer, Livien, Zhong, Zeyun, Martin, Manuel, Diederichs, Frederik, Stiefelhagen, Rainer |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset
by: Lerch, David J., et al.
Published: (2026)
by: Lerch, David J., et al.
Published: (2026)
QueryMamba: A Mamba-Based Encoder-Decoder Architecture with a Statistical Verb-Noun Interaction Module for Video Action Forecasting @ Ego4D Long-Term Action Anticipation Challenge 2024
by: Zhong, Zeyun, et al.
Published: (2024)
by: Zhong, Zeyun, et al.
Published: (2024)
FlowNar: Scalable Streaming Narration for Long-Form Videos
by: Zhong, Zeyun, et al.
Published: (2026)
by: Zhong, Zeyun, et al.
Published: (2026)
OmniFall: From Staged Through Synthetic to Wild, A Unified Multi-Domain Dataset for Robust Fall Detection
by: Schneider, David, et al.
Published: (2025)
by: Schneider, David, et al.
Published: (2025)
FakeOut: Leveraging Out-of-domain Self-supervision for Multi-modal Video Deepfake Detection
by: Knafo, Gil, et al.
Published: (2022)
by: Knafo, Gil, et al.
Published: (2022)
EyeCue: Driver Cognitive Distraction Detection via Gaze-Empowered Egocentric Video Understanding
by: Zhang, Lang, et al.
Published: (2026)
by: Zhang, Lang, et al.
Published: (2026)
ViT-DD: Multi-Task Vision Transformer for Semi-Supervised Driver Distraction Detection
by: Ma, Yunsheng, et al.
Published: (2022)
by: Ma, Yunsheng, et al.
Published: (2022)
A Contextual Analysis of Driver-Facing and Dual-View Video Inputs for Distraction Detection in Naturalistic Driving Environments
by: Dontoh, Anthony, et al.
Published: (2025)
by: Dontoh, Anthony, et al.
Published: (2025)
mmWalk: Towards Multi-modal Multi-view Walking Assistance
by: Ying, Kedi, et al.
Published: (2025)
by: Ying, Kedi, et al.
Published: (2025)
Vision-Language Models can Identify Distracted Driver Behavior from Naturalistic Videos
by: Hasan, Md Zahid, et al.
Published: (2023)
by: Hasan, Md Zahid, et al.
Published: (2023)
DSDFormer: An Innovative Transformer-Mamba Framework for Robust High-Precision Driver Distraction Identification
by: Chen, Junzhou, et al.
Published: (2024)
by: Chen, Junzhou, et al.
Published: (2024)
IMPACT-CYCLE: A Contract-Based Multi-Agent System for Claim-Level Supervisory Correction of Long-Video Semantic Memory
by: Kong, Weitong, et al.
Published: (2026)
by: Kong, Weitong, et al.
Published: (2026)
Towards Infusing Auxiliary Knowledge for Distracted Driver Detection
by: Balappanawar, Ishwar B, et al.
Published: (2024)
by: Balappanawar, Ishwar B, et al.
Published: (2024)
CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding
by: Liu, Yunze, et al.
Published: (2024)
by: Liu, Yunze, et al.
Published: (2024)
Enhancing Road Safety: Real-Time Detection of Driver Distraction through Convolutional Neural Networks
by: Sheikh, Amaan Aijaz, et al.
Published: (2024)
by: Sheikh, Amaan Aijaz, et al.
Published: (2024)
Rethinking Top Probability from Multi-view for Distracted Driver Behaviour Localization
by: Nguyen, Quang Vinh, et al.
Published: (2024)
by: Nguyen, Quang Vinh, et al.
Published: (2024)
Collaboratively Self-supervised Video Representation Learning for Action Recognition
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
A Self-supervised Motion Representation for Portrait Video Generation
by: Zhang, Qiyuan, et al.
Published: (2025)
by: Zhang, Qiyuan, et al.
Published: (2025)
Spiking-DD: Neuromorphic Event Camera based Driver Distraction Detection with Spiking Neural Network
by: Shariff, Waseem, et al.
Published: (2024)
by: Shariff, Waseem, et al.
Published: (2024)
Self-supervised 3D Patient Modeling with Multi-modal Attentive Fusion
by: Zheng, Meng, et al.
Published: (2024)
by: Zheng, Meng, et al.
Published: (2024)
A Survey on Deep Learning Techniques for Action Anticipation
by: Zhong, Zeyun, et al.
Published: (2023)
by: Zhong, Zeyun, et al.
Published: (2023)
Exploring Self-supervised Skeleton-based Action Recognition in Occluded Environments
by: Chen, Yifei, et al.
Published: (2023)
by: Chen, Yifei, et al.
Published: (2023)
Steering and Rectifying Latent Representation Manifolds in Frozen Multi-modal LLMs for Video Anomaly Detection
by: Cai, Zhaolin, et al.
Published: (2026)
by: Cai, Zhaolin, et al.
Published: (2026)
Predicting Local Climate Zones using Urban Morphometrics and Satellite Imagery
by: Majer, Hugo, et al.
Published: (2026)
by: Majer, Hugo, et al.
Published: (2026)
PQ-DAF: Pose-driven Quality-controlled Data Augmentation for Data-scarce Driver Distraction Detection
by: Sun, Haibin, et al.
Published: (2025)
by: Sun, Haibin, et al.
Published: (2025)
Analyzing Local Representations of Self-supervised Vision Transformers
by: Vanyan, Ani, et al.
Published: (2023)
by: Vanyan, Ani, et al.
Published: (2023)
Zero-Shot Distracted Driver Detection via Vision Language Models with Double Decoupling
by: Miyata, Takamichi, et al.
Published: (2026)
by: Miyata, Takamichi, et al.
Published: (2026)
Mars Traversability Prediction: A Multi-modal Self-supervised Approach for Costmap Generation
by: Xie, Zongwu, et al.
Published: (2025)
by: Xie, Zongwu, et al.
Published: (2025)
Self-supervised Video Instance Segmentation Can Boost Geographic Entity Alignment in Historical Maps
by: Xia, Xue, et al.
Published: (2024)
by: Xia, Xue, et al.
Published: (2024)
SCPNet: Unsupervised Cross-modal Homography Estimation via Intra-modal Self-supervised Learning
by: Zhang, Runmin, et al.
Published: (2024)
by: Zhang, Runmin, et al.
Published: (2024)
From Static to Dynamic: Exploring Self-supervised Image-to-Video Representation Transfer Learning
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
Car Damage Detection and Patch-to-Patch Self-supervised Image Alignment
by: Chen, Hanxiao
Published: (2024)
by: Chen, Hanxiao
Published: (2024)
Self-supervised Multi-actor Social Activity Understanding in Streaming Videos
by: Trehan, Shubham, et al.
Published: (2024)
by: Trehan, Shubham, et al.
Published: (2024)
Exploring Video-Based Driver Activity Recognition under Noisy Labels
by: Fan, Linjuan, et al.
Published: (2025)
by: Fan, Linjuan, et al.
Published: (2025)
MDMP: Multi-modal Diffusion for supervised Motion Predictions with uncertainty
by: Bringer, Leo, et al.
Published: (2024)
by: Bringer, Leo, et al.
Published: (2024)
When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation Learning
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Masked Diffusion as Self-supervised Representation Learner
by: Pan, Zixuan, et al.
Published: (2023)
by: Pan, Zixuan, et al.
Published: (2023)
Transformer-based Fusion of 2D-pose and Spatio-temporal Embeddings for Distracted Driver Action Recognition
by: Akdag, Erkut, et al.
Published: (2024)
by: Akdag, Erkut, et al.
Published: (2024)
MIFI: MultI-camera Feature Integration for Roust 3D Distracted Driver Activity Recognition
by: Kuang, Jian, et al.
Published: (2024)
by: Kuang, Jian, et al.
Published: (2024)
H-Flow: Self-supervised Human Scene Flow via Physics-inspired Joint Multi-modal Learning
by: Huang, Zhanbo, et al.
Published: (2026)
by: Huang, Zhanbo, et al.
Published: (2026)
Similar Items
-
Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset
by: Lerch, David J., et al.
Published: (2026) -
QueryMamba: A Mamba-Based Encoder-Decoder Architecture with a Statistical Verb-Noun Interaction Module for Video Action Forecasting @ Ego4D Long-Term Action Anticipation Challenge 2024
by: Zhong, Zeyun, et al.
Published: (2024) -
FlowNar: Scalable Streaming Narration for Long-Form Videos
by: Zhong, Zeyun, et al.
Published: (2026) -
OmniFall: From Staged Through Synthetic to Wild, A Unified Multi-Domain Dataset for Robust Fall Detection
by: Schneider, David, et al.
Published: (2025) -
FakeOut: Leveraging Out-of-domain Self-supervision for Multi-modal Video Deepfake Detection
by: Knafo, Gil, et al.
Published: (2022)