Robust Multiview Multimodal Driver Monitoring System Using Masked Multi-Head Self-Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Yiming, Sanchez, Victor, Nikan, Soodeh, Upadhyay, Devesh, Atote, Bhushan, Guha, Tanaya |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhanced Prototypical Part Network (EPPNet) For Explainable Image Classification Via Prototypes
by: Atote, Bhushan, et al.
Published: (2024)
by: Atote, Bhushan, et al.
Published: (2024)
CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise Classification
by: Ma, Yiming, et al.
Published: (2024)
by: Ma, Yiming, et al.
Published: (2024)
ZIP: Scalable Crowd Counting via Zero-Inflated Poisson Modeling
by: Ma, Yiming, et al.
Published: (2025)
by: Ma, Yiming, et al.
Published: (2025)
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning
by: Chen, Lihong, et al.
Published: (2025)
by: Chen, Lihong, et al.
Published: (2025)
TinyDrive: Multiscale Visual Question Answering with Selective Token Routing for Autonomous Driving
by: Hassani, Hossein, et al.
Published: (2025)
by: Hassani, Hossein, et al.
Published: (2025)
Interact with me: Joint Egocentric Forecasting of Intent to Interact, Attitude and Social Actions
by: Bian, Tongfei, et al.
Published: (2024)
by: Bian, Tongfei, et al.
Published: (2024)
Robust Understanding of Human-Robot Social Interactions through Multimodal Distillation
by: Bian, Tongfei, et al.
Published: (2025)
by: Bian, Tongfei, et al.
Published: (2025)
Efficient Transformer-based Hyper-parameter Optimization for Resource-constrained IoT Environments
by: Shaer, Ibrahim, et al.
Published: (2024)
by: Shaer, Ibrahim, et al.
Published: (2024)
Active Listener: Continuous Generation of Listener's Head Motion Response in Dyadic Interactions
by: Ghosh, Bishal, et al.
Published: (2024)
by: Ghosh, Bishal, et al.
Published: (2024)
Out-of-Distribution Detection with Attention Head Masking for Multimodal Document Classification
by: Constantinou, Christos, et al.
Published: (2024)
by: Constantinou, Christos, et al.
Published: (2024)
Structure is Supervision: Multiview Masked Autoencoders for Radiology
by: Laguna, Sonia, et al.
Published: (2025)
by: Laguna, Sonia, et al.
Published: (2025)
Improving Vision Transformers by Overlapping Heads in Multi-Head Self-Attention
by: Zhang, Tianxiao, et al.
Published: (2024)
by: Zhang, Tianxiao, et al.
Published: (2024)
Occlusion-aware Driver Monitoring System using the Driver Monitoring Dataset
by: Cañas, Paola Natalia, et al.
Published: (2025)
by: Cañas, Paola Natalia, et al.
Published: (2025)
DAOS: A Multimodal In-cabin Behavior Monitoring with Driver Action-Object Synergy Dataset
by: Li, Yiming, et al.
Published: (2026)
by: Li, Yiming, et al.
Published: (2026)
Interactive Multi-Head Self-Attention with Linear Complexity
by: Kang, Hankyul, et al.
Published: (2024)
by: Kang, Hankyul, et al.
Published: (2024)
Agentic Pipeline for Self-Synchronized Multiview Joint Angle Monitoring in Uncalibrated Environments
by: Yu, Juncheng, et al.
Published: (2026)
by: Yu, Juncheng, et al.
Published: (2026)
FROMAT: Multiview Material Appearance Transfer via Few-Shot Self-Attention Adaptation
by: Kompanowski, Hubert, et al.
Published: (2025)
by: Kompanowski, Hubert, et al.
Published: (2025)
Multi-layer Learnable Attention Mask for Multimodal Tasks
by: Barrios, Wayner, et al.
Published: (2024)
by: Barrios, Wayner, et al.
Published: (2024)
Resource-Efficient Multiview Perception: Integrating Semantic Masking with Masked Autoencoders
by: Dakic, Kosta, et al.
Published: (2024)
by: Dakic, Kosta, et al.
Published: (2024)
T-MASK: Temporal Masking for Probing Foundation Models across Camera Views in Driver Monitoring
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2025)
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2025)
Self-Supervised Multiview Xray Matching
by: Dabboussi, Mohamad, et al.
Published: (2025)
by: Dabboussi, Mohamad, et al.
Published: (2025)
Curvature-Regularized Variational Autoencoder for 3D Scene Reconstruction from Sparse Depth
by: Yousefi, Maryam, et al.
Published: (2025)
by: Yousefi, Maryam, et al.
Published: (2025)
MoH: Multi-Head Attention as Mixture-of-Head Attention
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
DeAR: Fine-Grained VLM Adaptation by Decomposing Attention Head Roles
by: Ma, Yiming, et al.
Published: (2026)
by: Ma, Yiming, et al.
Published: (2026)
Multimodal and Multiview Deep Fusion for Autonomous Marine Navigation
by: Dagdilelis, Dimitrios, et al.
Published: (2025)
by: Dagdilelis, Dimitrios, et al.
Published: (2025)
Low-Latency Embedded Driver Monitoring System with a Multi-Task Neural Network
by: Scribano, Carmelo, et al.
Published: (2026)
by: Scribano, Carmelo, et al.
Published: (2026)
Multiview Self-Representation Learning across Heterogeneous Views
by: Chen, Jie, et al.
Published: (2026)
by: Chen, Jie, et al.
Published: (2026)
Constrained Multiview Representation for Self-supervised Contrastive Learning
by: Dai, Siyuan, et al.
Published: (2024)
by: Dai, Siyuan, et al.
Published: (2024)
Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset
by: Lerch, David J., et al.
Published: (2026)
by: Lerch, David J., et al.
Published: (2026)
Exploration of VLMs for Driver Monitoring Systems Applications
by: Cañas, Paola Natalia, et al.
Published: (2025)
by: Cañas, Paola Natalia, et al.
Published: (2025)
VideoClusterNet: Self-Supervised and Adaptive Face Clustering For Videos
by: Walawalkar, Devesh, et al.
Published: (2024)
by: Walawalkar, Devesh, et al.
Published: (2024)
Multi-modal Video Representation Alignment for Robust Self-supervised Driver Distraction Detection
by: Lerch, David J., et al.
Published: (2026)
by: Lerch, David J., et al.
Published: (2026)
Identity Preserving 3D Head Stylization with Multiview Score Distillation
by: Bilecen, Bahri Batuhan, et al.
Published: (2024)
by: Bilecen, Bahri Batuhan, et al.
Published: (2024)
Sports Analysis and VR Viewing System Based on Player Tracking and Pose Estimation with Multimodal and Multiview Sensors
by: Guo, Wenxuan, et al.
Published: (2024)
by: Guo, Wenxuan, et al.
Published: (2024)
Enhanced Survival Prediction in Head and Neck Cancer Using Convolutional Block Attention and Multimodal Data Fusion
by: Farooq, Aiman, et al.
Published: (2024)
by: Farooq, Aiman, et al.
Published: (2024)
Learning Subspace-Preserving Sparse Attention Graphs from Heterogeneous Multiview Data
by: Chen, Jie, et al.
Published: (2026)
by: Chen, Jie, et al.
Published: (2026)
Alignment Scores: Robust Metrics for Multiview Pose Accuracy Evaluation
by: Lee, Seong Hun, et al.
Published: (2024)
by: Lee, Seong Hun, et al.
Published: (2024)
MEAT: Multiview Diffusion Model for Human Generation on Megapixels with Mesh Attention
by: Wang, Yuhan, et al.
Published: (2025)
by: Wang, Yuhan, et al.
Published: (2025)
Efficient Masked Image Compression with Position-Indexed Self-Attention
by: Dai, Chengjie, et al.
Published: (2025)
by: Dai, Chengjie, et al.
Published: (2025)
VLM-E2E: Enhancing End-to-End Autonomous Driving with Multimodal Driver Attention Fusion
by: Liu, Pei, et al.
Published: (2025)
by: Liu, Pei, et al.
Published: (2025)
Similar Items
-
Enhanced Prototypical Part Network (EPPNet) For Explainable Image Classification Via Prototypes
by: Atote, Bhushan, et al.
Published: (2024) -
CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise Classification
by: Ma, Yiming, et al.
Published: (2024) -
ZIP: Scalable Crowd Counting via Zero-Inflated Poisson Modeling
by: Ma, Yiming, et al.
Published: (2025) -
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning
by: Chen, Lihong, et al.
Published: (2025) -
TinyDrive: Multiscale Visual Question Answering with Selective Token Routing for Autonomous Driving
by: Hassani, Hossein, et al.
Published: (2025)