Exploring the Potential of Multi-Modal AI for Driving Hazard Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Charoenpitaks, Korawat, Nguyen, Van-Quang, Suganuma, Masanori, Takahashi, Masahiro, Niihara, Ryoma, Okatani, Takayuki |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TB-Bench: Training and Testing Multi-Modal AI for Understanding Spatio-Temporal Traffic Behaviors from Dashcam Images/Videos
by: Charoenpitaks, Korawat, et al.
Published: (2025)
by: Charoenpitaks, Korawat, et al.
Published: (2025)
Cascaded Multi-Scale Attention for Enhanced Multi-Scale Feature Extraction and Interaction with Low-Resolution Images
by: Lu, Xiangyong, et al.
Published: (2024)
by: Lu, Xiangyong, et al.
Published: (2024)
RefVSR++: Exploiting Reference Inputs for Reference-based Video Super-resolution
by: Zou, Han, et al.
Published: (2023)
by: Zou, Han, et al.
Published: (2023)
An Improved Method for Personalizing Diffusion Models
by: Zeng, Yan, et al.
Published: (2024)
by: Zeng, Yan, et al.
Published: (2024)
Open-vocabulary vs. Closed-set: Best Practice for Few-shot Object Detection Considering Text Describability
by: Hosoya, Yusuke, et al.
Published: (2024)
by: Hosoya, Yusuke, et al.
Published: (2024)
Rethinking Annotation for Object Detection: Is Annotating Small-size Instances Worth Its Cost?
by: Hosoya, Yusuke, et al.
Published: (2024)
by: Hosoya, Yusuke, et al.
Published: (2024)
Rethinking Open-Set Object Detection: Issues, a New Formulation, and Taxonomy
by: Hosoya, Yusuke, et al.
Published: (2022)
by: Hosoya, Yusuke, et al.
Published: (2022)
Rethinking Unsupervised Domain Adaptation for Semantic Segmentation
by: Wang, Zhijie, et al.
Published: (2022)
by: Wang, Zhijie, et al.
Published: (2022)
MS-DPPs: Multi-Source Determinantal Point Processes for Contextual Diversity Refinement of Composite Attributes in Text to Image Retrieval
by: Sogi, Naoya, et al.
Published: (2025)
by: Sogi, Naoya, et al.
Published: (2025)
RP-SLAM: Real-time Photorealistic SLAM with Efficient 3D Gaussian Splatting
by: Bai, Lizhi, et al.
Published: (2024)
by: Bai, Lizhi, et al.
Published: (2024)
360° Image Perception with MLLMs: A Comprehensive Benchmark and a Training-Free Method
by: Tran, Huyen T. T., et al.
Published: (2026)
by: Tran, Huyen T. T., et al.
Published: (2026)
LAMS-Edit: Latent and Attention Mixing with Schedulers for Improved Content Preservation in Diffusion-Based Image and Style Editing
by: Fu, Wingwa, et al.
Published: (2026)
by: Fu, Wingwa, et al.
Published: (2026)
Temporal Insight Enhancement: Mitigating Temporal Hallucination in Multimodal Large Language Models
by: Sun, Li, et al.
Published: (2024)
by: Sun, Li, et al.
Published: (2024)
Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments
by: Nguyen, Van Quang
Published: (2026)
by: Nguyen, Van Quang
Published: (2026)
TwinLiteNet+: An Enhanced Multi-Task Segmentation Model for Autonomous Driving
by: Che, Quang-Huy, et al.
Published: (2024)
by: Che, Quang-Huy, et al.
Published: (2024)
Guidance-base Diffusion Models for Improving Photoacoustic Image Quality
by: Eguchi, Tatsuhiro, et al.
Published: (2025)
by: Eguchi, Tatsuhiro, et al.
Published: (2025)
Lightweight Temporal Transformer Decomposition for Federated Autonomous Driving
by: Do, Tuong, et al.
Published: (2025)
by: Do, Tuong, et al.
Published: (2025)
Action-Agnostic Point-Level Supervision for Temporal Action Detection
by: Yoshida, Shuhei M., et al.
Published: (2024)
by: Yoshida, Shuhei M., et al.
Published: (2024)
Label-Efficient Cross-Modality Generalization for Liver Segmentation in Multi-Phase MRI
by: Bui-Tran, Quang-Khai, et al.
Published: (2025)
by: Bui-Tran, Quang-Khai, et al.
Published: (2025)
ViOCRVQA: Novel Benchmark Dataset and Vision Reader for Visual Question Answering by Understanding Vietnamese Text in Images
by: Pham, Huy Quang, et al.
Published: (2024)
by: Pham, Huy Quang, et al.
Published: (2024)
DeepInteraction++: Multi-Modality Interaction for Autonomous Driving
by: Yang, Zeyu, et al.
Published: (2024)
by: Yang, Zeyu, et al.
Published: (2024)
TF-SASM: Training-free Spatial-aware Sparse Memory for Multi-object Tracking
by: Nguyen-Quang, Thuc, et al.
Published: (2024)
by: Nguyen-Quang, Thuc, et al.
Published: (2024)
RealDriveSim: A Realistic Multi-Modal Multi-Task Synthetic Dataset for Autonomous Driving
by: Jadon, Arpit, et al.
Published: (2025)
by: Jadon, Arpit, et al.
Published: (2025)
CoReTab: Improving Multimodal Table Understanding with Code-driven Reasoning
by: Nguyen, Van-Quang, et al.
Published: (2026)
by: Nguyen, Van-Quang, et al.
Published: (2026)
123D: Unifying Multi-Modal Autonomous Driving Data at Scale
by: Dauner, Daniel, et al.
Published: (2026)
by: Dauner, Daniel, et al.
Published: (2026)
General Hazard Detection
by: Ng, Stephanie, et al.
Published: (2026)
by: Ng, Stephanie, et al.
Published: (2026)
FurniMAS: Language-Guided Furniture Decoration using Multi-Agent System
by: Nguyen, Toan, et al.
Published: (2025)
by: Nguyen, Toan, et al.
Published: (2025)
UIT-DarkCow team at ImageCLEFmedical Caption 2024: Diagnostic Captioning for Radiology Images Efficiency with Transformer Models
by: Van Nguyen, Quan, et al.
Published: (2024)
by: Van Nguyen, Quan, et al.
Published: (2024)
DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving
by: Cui, Erfei, et al.
Published: (2023)
by: Cui, Erfei, et al.
Published: (2023)
Modality-Specific Enhancement and Complementary Fusion for Semi-Supervised Multi-Modal Brain Tumor Segmentation
by: Chung, Tien-Dat, et al.
Published: (2025)
by: Chung, Tien-Dat, et al.
Published: (2025)
Exploring Modality-Aware Fusion and Decoupled Temporal Propagation for Multi-Modal Object Tracking
by: Wang, Shilei, et al.
Published: (2026)
by: Wang, Shilei, et al.
Published: (2026)
CooperDrive: Enhancing Driving Decisions Through Cooperative Perception
by: Qu, Deyuan, et al.
Published: (2026)
by: Qu, Deyuan, et al.
Published: (2026)
Zero-shot Hazard Identification in Autonomous Driving: A Case Study on the COOOL Benchmark
by: Picek, Lukas, et al.
Published: (2024)
by: Picek, Lukas, et al.
Published: (2024)
HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving
by: Wu, Zehuan, et al.
Published: (2024)
by: Wu, Zehuan, et al.
Published: (2024)
Graph-Based Multi-Modal Sensor Fusion for Autonomous Driving
by: Sani, Depanshu, et al.
Published: (2024)
by: Sani, Depanshu, et al.
Published: (2024)
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition
by: Nguyen, Thai-Binh, et al.
Published: (2025)
by: Nguyen, Thai-Binh, et al.
Published: (2025)
Exploring Generative AI for Sim2Real in Driving Data Synthesis
by: Zhao, Haonan, et al.
Published: (2024)
by: Zhao, Haonan, et al.
Published: (2024)
MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer
by: Wu, Guile, et al.
Published: (2025)
by: Wu, Guile, et al.
Published: (2025)
KTVIC: A Vietnamese Image Captioning Dataset on the Life Domain
by: Pham, Anh-Cuong, et al.
Published: (2024)
by: Pham, Anh-Cuong, et al.
Published: (2024)
MissBench: Benchmarking Multimodal Affective Analysis under Imbalanced Missing Modalities
by: Pham, Tien Anh, et al.
Published: (2026)
by: Pham, Tien Anh, et al.
Published: (2026)
Similar Items
-
TB-Bench: Training and Testing Multi-Modal AI for Understanding Spatio-Temporal Traffic Behaviors from Dashcam Images/Videos
by: Charoenpitaks, Korawat, et al.
Published: (2025) -
Cascaded Multi-Scale Attention for Enhanced Multi-Scale Feature Extraction and Interaction with Low-Resolution Images
by: Lu, Xiangyong, et al.
Published: (2024) -
RefVSR++: Exploiting Reference Inputs for Reference-based Video Super-resolution
by: Zou, Han, et al.
Published: (2023) -
An Improved Method for Personalizing Diffusion Models
by: Zeng, Yan, et al.
Published: (2024) -
Open-vocabulary vs. Closed-set: Best Practice for Few-shot Object Detection Considering Text Describability
by: Hosoya, Yusuke, et al.
Published: (2024)