Robust Multiview Multimodal Driver Monitoring System Using Masked Multi-Head Self-Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Yiming, Sanchez, Victor, Nikan, Soodeh, Upadhyay, Devesh, Atote, Bhushan, Guha, Tanaya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhanced Prototypical Part Network (EPPNet) For Explainable Image Classification Via Prototypes
von: Atote, Bhushan, et al.
Veröffentlicht: (2024)
von: Atote, Bhushan, et al.
Veröffentlicht: (2024)
CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise Classification
von: Ma, Yiming, et al.
Veröffentlicht: (2024)
von: Ma, Yiming, et al.
Veröffentlicht: (2024)
ZIP: Scalable Crowd Counting via Zero-Inflated Poisson Modeling
von: Ma, Yiming, et al.
Veröffentlicht: (2025)
von: Ma, Yiming, et al.
Veröffentlicht: (2025)
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning
von: Chen, Lihong, et al.
Veröffentlicht: (2025)
von: Chen, Lihong, et al.
Veröffentlicht: (2025)
TinyDrive: Multiscale Visual Question Answering with Selective Token Routing for Autonomous Driving
von: Hassani, Hossein, et al.
Veröffentlicht: (2025)
von: Hassani, Hossein, et al.
Veröffentlicht: (2025)
Interact with me: Joint Egocentric Forecasting of Intent to Interact, Attitude and Social Actions
von: Bian, Tongfei, et al.
Veröffentlicht: (2024)
von: Bian, Tongfei, et al.
Veröffentlicht: (2024)
Robust Understanding of Human-Robot Social Interactions through Multimodal Distillation
von: Bian, Tongfei, et al.
Veröffentlicht: (2025)
von: Bian, Tongfei, et al.
Veröffentlicht: (2025)
Efficient Transformer-based Hyper-parameter Optimization for Resource-constrained IoT Environments
von: Shaer, Ibrahim, et al.
Veröffentlicht: (2024)
von: Shaer, Ibrahim, et al.
Veröffentlicht: (2024)
Active Listener: Continuous Generation of Listener's Head Motion Response in Dyadic Interactions
von: Ghosh, Bishal, et al.
Veröffentlicht: (2024)
von: Ghosh, Bishal, et al.
Veröffentlicht: (2024)
Out-of-Distribution Detection with Attention Head Masking for Multimodal Document Classification
von: Constantinou, Christos, et al.
Veröffentlicht: (2024)
von: Constantinou, Christos, et al.
Veröffentlicht: (2024)
Structure is Supervision: Multiview Masked Autoencoders for Radiology
von: Laguna, Sonia, et al.
Veröffentlicht: (2025)
von: Laguna, Sonia, et al.
Veröffentlicht: (2025)
Improving Vision Transformers by Overlapping Heads in Multi-Head Self-Attention
von: Zhang, Tianxiao, et al.
Veröffentlicht: (2024)
von: Zhang, Tianxiao, et al.
Veröffentlicht: (2024)
Occlusion-aware Driver Monitoring System using the Driver Monitoring Dataset
von: Cañas, Paola Natalia, et al.
Veröffentlicht: (2025)
von: Cañas, Paola Natalia, et al.
Veröffentlicht: (2025)
DAOS: A Multimodal In-cabin Behavior Monitoring with Driver Action-Object Synergy Dataset
von: Li, Yiming, et al.
Veröffentlicht: (2026)
von: Li, Yiming, et al.
Veröffentlicht: (2026)
Interactive Multi-Head Self-Attention with Linear Complexity
von: Kang, Hankyul, et al.
Veröffentlicht: (2024)
von: Kang, Hankyul, et al.
Veröffentlicht: (2024)
Agentic Pipeline for Self-Synchronized Multiview Joint Angle Monitoring in Uncalibrated Environments
von: Yu, Juncheng, et al.
Veröffentlicht: (2026)
von: Yu, Juncheng, et al.
Veröffentlicht: (2026)
FROMAT: Multiview Material Appearance Transfer via Few-Shot Self-Attention Adaptation
von: Kompanowski, Hubert, et al.
Veröffentlicht: (2025)
von: Kompanowski, Hubert, et al.
Veröffentlicht: (2025)
Multi-layer Learnable Attention Mask for Multimodal Tasks
von: Barrios, Wayner, et al.
Veröffentlicht: (2024)
von: Barrios, Wayner, et al.
Veröffentlicht: (2024)
Resource-Efficient Multiview Perception: Integrating Semantic Masking with Masked Autoencoders
von: Dakic, Kosta, et al.
Veröffentlicht: (2024)
von: Dakic, Kosta, et al.
Veröffentlicht: (2024)
T-MASK: Temporal Masking for Probing Foundation Models across Camera Views in Driver Monitoring
von: Ponbagavathi, Thinesh Thiyakesan, et al.
Veröffentlicht: (2025)
von: Ponbagavathi, Thinesh Thiyakesan, et al.
Veröffentlicht: (2025)
Self-Supervised Multiview Xray Matching
von: Dabboussi, Mohamad, et al.
Veröffentlicht: (2025)
von: Dabboussi, Mohamad, et al.
Veröffentlicht: (2025)
Curvature-Regularized Variational Autoencoder for 3D Scene Reconstruction from Sparse Depth
von: Yousefi, Maryam, et al.
Veröffentlicht: (2025)
von: Yousefi, Maryam, et al.
Veröffentlicht: (2025)
MoH: Multi-Head Attention as Mixture-of-Head Attention
von: Jin, Peng, et al.
Veröffentlicht: (2024)
von: Jin, Peng, et al.
Veröffentlicht: (2024)
DeAR: Fine-Grained VLM Adaptation by Decomposing Attention Head Roles
von: Ma, Yiming, et al.
Veröffentlicht: (2026)
von: Ma, Yiming, et al.
Veröffentlicht: (2026)
Multimodal and Multiview Deep Fusion for Autonomous Marine Navigation
von: Dagdilelis, Dimitrios, et al.
Veröffentlicht: (2025)
von: Dagdilelis, Dimitrios, et al.
Veröffentlicht: (2025)
Low-Latency Embedded Driver Monitoring System with a Multi-Task Neural Network
von: Scribano, Carmelo, et al.
Veröffentlicht: (2026)
von: Scribano, Carmelo, et al.
Veröffentlicht: (2026)
Multiview Self-Representation Learning across Heterogeneous Views
von: Chen, Jie, et al.
Veröffentlicht: (2026)
von: Chen, Jie, et al.
Veröffentlicht: (2026)
Constrained Multiview Representation for Self-supervised Contrastive Learning
von: Dai, Siyuan, et al.
Veröffentlicht: (2024)
von: Dai, Siyuan, et al.
Veröffentlicht: (2024)
Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset
von: Lerch, David J., et al.
Veröffentlicht: (2026)
von: Lerch, David J., et al.
Veröffentlicht: (2026)
Exploration of VLMs for Driver Monitoring Systems Applications
von: Cañas, Paola Natalia, et al.
Veröffentlicht: (2025)
von: Cañas, Paola Natalia, et al.
Veröffentlicht: (2025)
VideoClusterNet: Self-Supervised and Adaptive Face Clustering For Videos
von: Walawalkar, Devesh, et al.
Veröffentlicht: (2024)
von: Walawalkar, Devesh, et al.
Veröffentlicht: (2024)
Multi-modal Video Representation Alignment for Robust Self-supervised Driver Distraction Detection
von: Lerch, David J., et al.
Veröffentlicht: (2026)
von: Lerch, David J., et al.
Veröffentlicht: (2026)
Identity Preserving 3D Head Stylization with Multiview Score Distillation
von: Bilecen, Bahri Batuhan, et al.
Veröffentlicht: (2024)
von: Bilecen, Bahri Batuhan, et al.
Veröffentlicht: (2024)
Sports Analysis and VR Viewing System Based on Player Tracking and Pose Estimation with Multimodal and Multiview Sensors
von: Guo, Wenxuan, et al.
Veröffentlicht: (2024)
von: Guo, Wenxuan, et al.
Veröffentlicht: (2024)
Enhanced Survival Prediction in Head and Neck Cancer Using Convolutional Block Attention and Multimodal Data Fusion
von: Farooq, Aiman, et al.
Veröffentlicht: (2024)
von: Farooq, Aiman, et al.
Veröffentlicht: (2024)
Learning Subspace-Preserving Sparse Attention Graphs from Heterogeneous Multiview Data
von: Chen, Jie, et al.
Veröffentlicht: (2026)
von: Chen, Jie, et al.
Veröffentlicht: (2026)
Alignment Scores: Robust Metrics for Multiview Pose Accuracy Evaluation
von: Lee, Seong Hun, et al.
Veröffentlicht: (2024)
von: Lee, Seong Hun, et al.
Veröffentlicht: (2024)
MEAT: Multiview Diffusion Model for Human Generation on Megapixels with Mesh Attention
von: Wang, Yuhan, et al.
Veröffentlicht: (2025)
von: Wang, Yuhan, et al.
Veröffentlicht: (2025)
Efficient Masked Image Compression with Position-Indexed Self-Attention
von: Dai, Chengjie, et al.
Veröffentlicht: (2025)
von: Dai, Chengjie, et al.
Veröffentlicht: (2025)
VLM-E2E: Enhancing End-to-End Autonomous Driving with Multimodal Driver Attention Fusion
von: Liu, Pei, et al.
Veröffentlicht: (2025)
von: Liu, Pei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Enhanced Prototypical Part Network (EPPNet) For Explainable Image Classification Via Prototypes
von: Atote, Bhushan, et al.
Veröffentlicht: (2024) -
CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise Classification
von: Ma, Yiming, et al.
Veröffentlicht: (2024) -
ZIP: Scalable Crowd Counting via Zero-Inflated Poisson Modeling
von: Ma, Yiming, et al.
Veröffentlicht: (2025) -
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning
von: Chen, Lihong, et al.
Veröffentlicht: (2025) -
TinyDrive: Multiscale Visual Question Answering with Selective Token Routing for Autonomous Driving
von: Hassani, Hossein, et al.
Veröffentlicht: (2025)