DAOS: A Multimodal In-cabin Behavior Monitoring with Driver Action-Object Synergy Dataset
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yiming, Cai, Chen, Liu, Tianyi, Lin, Dan, Wang, Wenqian, Liang, Wenfei, Li, Bingbing, Yap, Kim-Hui |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MultiFuser: Multimodal Fusion Transformer for Enhanced Driver Action Recognition
by: Wang, Ruoyu, et al.
Published: (2024)
by: Wang, Ruoyu, et al.
Published: (2024)
Mixture-of-Modality-Experts with Holistic Token Learning for Fine-Grained Multimodal Visual Analytics in Driver Action Recognition
by: Liu, Tianyi, et al.
Published: (2026)
by: Liu, Tianyi, et al.
Published: (2026)
Open World Object Detection: A Survey
by: Li, Yiming, et al.
Published: (2024)
by: Li, Yiming, et al.
Published: (2024)
Multi-modality action recognition based on dual feature shift in vehicle cabin monitoring
by: Lin, Dan, et al.
Published: (2024)
by: Lin, Dan, et al.
Published: (2024)
CM2-Net: Continual Cross-Modal Mapping Network for Driver Action Recognition
by: Wang, Ruoyu, et al.
Published: (2024)
by: Wang, Ruoyu, et al.
Published: (2024)
Occlusion-aware Driver Monitoring System using the Driver Monitoring Dataset
by: Cañas, Paola Natalia, et al.
Published: (2025)
by: Cañas, Paola Natalia, et al.
Published: (2025)
Robust Multiview Multimodal Driver Monitoring System Using Masked Multi-Head Self-Attention
by: Ma, Yiming, et al.
Published: (2023)
by: Ma, Yiming, et al.
Published: (2023)
PhysDrive: A Multimodal Remote Physiological Measurement Dataset for In-vehicle Driver Monitoring
by: Wang, Jiyao, et al.
Published: (2025)
by: Wang, Jiyao, et al.
Published: (2025)
Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset
by: Lerch, David J., et al.
Published: (2026)
by: Lerch, David J., et al.
Published: (2026)
EraW-Net: Enhance-Refine-Align W-Net for Scene-Associated Driver Attention Estimation
by: Zhou, Jun, et al.
Published: (2024)
by: Zhou, Jun, et al.
Published: (2024)
HabitAction: A Video Dataset for Human Habitual Behavior Recognition
by: Li, Hongwu, et al.
Published: (2024)
by: Li, Hongwu, et al.
Published: (2024)
OpenDriver: An Open-Road Driver State Detection Dataset
by: Liu, Delong, et al.
Published: (2023)
by: Liu, Delong, et al.
Published: (2023)
MM-CamObj: A Comprehensive Multimodal Dataset for Camouflaged Object Scenarios
by: Ruan, Jiacheng, et al.
Published: (2024)
by: Ruan, Jiacheng, et al.
Published: (2024)
Instance-Adaptive and Geometric-Aware Keypoint Learning for Category-Level 6D Object Pose Estimation
by: Lin, Xiao, et al.
Published: (2024)
by: Lin, Xiao, et al.
Published: (2024)
MDPE: A Multimodal Deception Dataset with Personality and Emotional Characteristics
by: Cai, Cong, et al.
Published: (2024)
by: Cai, Cong, et al.
Published: (2024)
MuLan: Multimodal-LLM Agent for Progressive and Interactive Multi-Object Diffusion
by: Li, Sen, et al.
Published: (2024)
by: Li, Sen, et al.
Published: (2024)
SmartWilds: Multimodal Wildlife Monitoring Dataset
by: Kline, Jenna, et al.
Published: (2025)
by: Kline, Jenna, et al.
Published: (2025)
PDB: Not All Drivers Are the Same -- A Personalized Dataset for Understanding Driving Behavior
by: Wei, Chuheng, et al.
Published: (2025)
by: Wei, Chuheng, et al.
Published: (2025)
STDA: Spatio-Temporal Dual-Encoder Network Incorporating Driver Attention to Predict Driver Behaviors Under Safety-Critical Scenarios
by: Xu, Dongyang, et al.
Published: (2024)
by: Xu, Dongyang, et al.
Published: (2024)
Multiagent Multitraversal Multimodal Self-Driving: Open MARS Dataset
by: Li, Yiming, et al.
Published: (2024)
by: Li, Yiming, et al.
Published: (2024)
MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection
by: Hu, Mengxue, et al.
Published: (2025)
by: Hu, Mengxue, et al.
Published: (2025)
ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition
by: Zhou, Jiaming, et al.
Published: (2024)
by: Zhou, Jiaming, et al.
Published: (2024)
Towards Small Object Editing: A Benchmark Dataset and A Training-Free Approach
by: Pan, Qihe, et al.
Published: (2024)
by: Pan, Qihe, et al.
Published: (2024)
SRasP: Self-Reorientation Adversarial Style Perturbation for Cross-Domain Few-Shot Learning
by: Li, Wenqian, et al.
Published: (2026)
by: Li, Wenqian, et al.
Published: (2026)
SVasP: Self-Versatility Adversarial Style Perturbation for Cross-Domain Few-Shot Learning
by: Li, Wenqian, et al.
Published: (2024)
by: Li, Wenqian, et al.
Published: (2024)
Delving into Cascaded Instability: A Lipschitz Continuity View on Image Restoration and Object Detection Synergy
by: Zhao, Qing, et al.
Published: (2025)
by: Zhao, Qing, et al.
Published: (2025)
CL-HOI: Cross-Level Human-Object Interaction Distillation from Vision Large Language Models
by: Gao, Jianjun, et al.
Published: (2024)
by: Gao, Jianjun, et al.
Published: (2024)
Efficient Temporal Sentence Grounding in Videos with Multi-Teacher Knowledge Distillation
by: Liang, Renjie, et al.
Published: (2023)
by: Liang, Renjie, et al.
Published: (2023)
3DReflecNet: A Large-Scale Dataset for 3D Reconstruction of Reflective, Transparent, and Low-Texture Objects
by: Liang, Zhicheng, et al.
Published: (2026)
by: Liang, Zhicheng, et al.
Published: (2026)
One Shot is Enough for Sequential Infrared Small Target Segmentation
by: Dan, Bingbing, et al.
Published: (2024)
by: Dan, Bingbing, et al.
Published: (2024)
TDS-CLIP: Temporal Difference Side Network for Efficient VideoAction Recognition
by: Wang, Bin, et al.
Published: (2024)
by: Wang, Bin, et al.
Published: (2024)
EventCrab: Harnessing Frame and Point Synergy for Event-based Action Recognition and Beyond
by: Cao, Meiqi, et al.
Published: (2024)
by: Cao, Meiqi, et al.
Published: (2024)
Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
by: Liu, Tianyi, et al.
Published: (2025)
by: Liu, Tianyi, et al.
Published: (2025)
CVBench: Benchmarking Cross-Video Synergies for Complex Multimodal Reasoning
by: Zhu, Nannan, et al.
Published: (2025)
by: Zhu, Nannan, et al.
Published: (2025)
YOLO-DS: Fine-Grained Feature Decoupling via Dual-Statistic Synergy Operator for Object Detection
by: Huang, Lin, et al.
Published: (2026)
by: Huang, Lin, et al.
Published: (2026)
PersonaX: Multimodal Datasets with LLM-Inferred Behavior Traits
by: Li, Loka, et al.
Published: (2025)
by: Li, Loka, et al.
Published: (2025)
MindDriver: Introducing Progressive Multimodal Reasoning for Autonomous Driving
by: Zhang, Lingjun, et al.
Published: (2026)
by: Zhang, Lingjun, et al.
Published: (2026)
Accelerating Diffusion-based Video Editing via Heterogeneous Caching: Beyond Full Computing at Sampled Denoising Timestep
by: Liu, Tianyi, et al.
Published: (2026)
by: Liu, Tianyi, et al.
Published: (2026)
A Multimodal Dataset for Enhancing Industrial Task Monitoring and Engagement Prediction
by: Mehta, Naval Kishore, et al.
Published: (2025)
by: Mehta, Naval Kishore, et al.
Published: (2025)
MapGlue: Multimodal Remote Sensing Image Matching
by: Wu, Peihao, et al.
Published: (2025)
by: Wu, Peihao, et al.
Published: (2025)
Similar Items
-
MultiFuser: Multimodal Fusion Transformer for Enhanced Driver Action Recognition
by: Wang, Ruoyu, et al.
Published: (2024) -
Mixture-of-Modality-Experts with Holistic Token Learning for Fine-Grained Multimodal Visual Analytics in Driver Action Recognition
by: Liu, Tianyi, et al.
Published: (2026) -
Open World Object Detection: A Survey
by: Li, Yiming, et al.
Published: (2024) -
Multi-modality action recognition based on dual feature shift in vehicle cabin monitoring
by: Lin, Dan, et al.
Published: (2024) -
CM2-Net: Continual Cross-Modal Mapping Network for Driver Action Recognition
by: Wang, Ruoyu, et al.
Published: (2024)