Multi-modal Sensor Fusion for Auto Driving Perception: A Survey
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Keli, Shi, Botian, Li, Xiang, Li, Xin, Huang, Siyuan, Li, Yikang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Survey on Dynamic Neural Networks: from Computer Vision to Multi-modal Sensor Fusion
von: Montello, Fabio, et al.
Veröffentlicht: (2025)
von: Montello, Fabio, et al.
Veröffentlicht: (2025)
3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding
von: Chen, Yiping, et al.
Veröffentlicht: (2026)
von: Chen, Yiping, et al.
Veröffentlicht: (2026)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
von: Li, Huibin, et al.
Veröffentlicht: (2025)
von: Li, Huibin, et al.
Veröffentlicht: (2025)
VisRL: Intention-Driven Visual Perception via Reinforced Reasoning
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025)
4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026)
CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment
von: Kang, Li, et al.
Veröffentlicht: (2026)
von: Kang, Li, et al.
Veröffentlicht: (2026)
Infrastructure-Centric World Models: Bridging Temporal Depth and Spatial Breadth for Roadside Perception
von: Meng, Siyuan, et al.
Veröffentlicht: (2026)
von: Meng, Siyuan, et al.
Veröffentlicht: (2026)
Hierarchical Feature-level Reverse Propagation for Post-Training Neural Networks
von: Ding, Ni, et al.
Veröffentlicht: (2025)
von: Ding, Ni, et al.
Veröffentlicht: (2025)
Enhancing Road Crack Detection Accuracy with BsS-YOLO: Optimizing Feature Fusion and Attention Mechanisms
von: Tang, Jiaze, et al.
Veröffentlicht: (2024)
von: Tang, Jiaze, et al.
Veröffentlicht: (2024)
Context-Aware Network Based on Multi-scale Spatio-temporal Attention for Action Recognition in Videos
von: Li, Xiaoyang, et al.
Veröffentlicht: (2025)
von: Li, Xiaoyang, et al.
Veröffentlicht: (2025)
One Leaf Reveals the Season: Occlusion-Based Contrastive Learning with Semantic-Aware Views for Efficient Visual Representation
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)
GMAC: Global Multi-View Constraint for Automatic Multi-Camera Extrinsic Calibration
von: Sun, Chentian
Veröffentlicht: (2026)
von: Sun, Chentian
Veröffentlicht: (2026)
Distinguishing Visually Similar Actions: Prompt-Guided Semantic Prototype Modulation for Few-Shot Action Recognition
von: Li, Xiaoyang, et al.
Veröffentlicht: (2025)
von: Li, Xiaoyang, et al.
Veröffentlicht: (2025)
Flexible-weighted Chamfer Distance: Enhanced Objective Function for Point Cloud Completion
von: Li, Jie, et al.
Veröffentlicht: (2025)
von: Li, Jie, et al.
Veröffentlicht: (2025)
Multi-modal Loop Closure Detection with Foundation Models in Severely Unstructured Environments
von: Gonzalez, Laura Alejandra Encinar, et al.
Veröffentlicht: (2025)
von: Gonzalez, Laura Alejandra Encinar, et al.
Veröffentlicht: (2025)
FocusedAD: Character-centric Movie Audio Description
von: Ye, Xiaojun, et al.
Veröffentlicht: (2025)
von: Ye, Xiaojun, et al.
Veröffentlicht: (2025)
A Challenging Benchmark of Anime Style Recognition
von: Li, Haotang, et al.
Veröffentlicht: (2022)
von: Li, Haotang, et al.
Veröffentlicht: (2022)
Unified Auto-Encoding with Masked Diffusion
von: Hansen-Estruch, Philippe, et al.
Veröffentlicht: (2024)
von: Hansen-Estruch, Philippe, et al.
Veröffentlicht: (2024)
OCC-MLLM-CoT-Alpha: Towards Multi-stage Occlusion Recognition Based on Large Language Models via 3D-Aware Supervision and Chain-of-Thoughts Guidance
von: Wang, Chaoyi, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyi, et al.
Veröffentlicht: (2025)
MSTA3D: Multi-scale Twin-attention for 3D Instance Segmentation
von: Tran, Duc Dang Trung, et al.
Veröffentlicht: (2024)
von: Tran, Duc Dang Trung, et al.
Veröffentlicht: (2024)
SPARK: Scalable Real-Time Point Cloud Aggregation with Multi-View Self-Calibration
von: Sun, Chentian
Veröffentlicht: (2026)
von: Sun, Chentian
Veröffentlicht: (2026)
FUSE-Flow: Scalable Real-Time Multi-View Point Cloud Reconstruction Using Confidence
von: Sun, Chentian
Veröffentlicht: (2026)
von: Sun, Chentian
Veröffentlicht: (2026)
FoR-Net: Learning to Focus on Hard Regions for Efficient Semantic Segmentation
von: Chan, Sheng-Wei, et al.
Veröffentlicht: (2026)
von: Chan, Sheng-Wei, et al.
Veröffentlicht: (2026)
Positive Style Accumulation: A Style Screening and Continuous Utilization Framework for Federated DG-ReID
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
Towards a Generalizable Fusion Architecture for Multimodal Object Detection
von: Berjawi, Jad, et al.
Veröffentlicht: (2025)
von: Berjawi, Jad, et al.
Veröffentlicht: (2025)
M3CAD: Towards Generic Cooperative Autonomous Driving Benchmark
von: Zhu, Morui, et al.
Veröffentlicht: (2025)
von: Zhu, Morui, et al.
Veröffentlicht: (2025)
RDPO: Real Data Preference Optimization for Physics Consistency Video Generation
von: Qian, Wenxu, et al.
Veröffentlicht: (2025)
von: Qian, Wenxu, et al.
Veröffentlicht: (2025)
WoundNet-Ensemble: A Novel IoMT System Integrating Self-Supervised Deep Learning and Multi-Model Fusion for Automated, High-Accuracy Wound Classification and Healing Progression Monitoring
von: Kiprono, Moses
Veröffentlicht: (2025)
von: Kiprono, Moses
Veröffentlicht: (2025)
Learning Association via Track-Detection Matching for Multi-Object Tracking
von: Adžemović, Momir
Veröffentlicht: (2025)
von: Adžemović, Momir
Veröffentlicht: (2025)
Auto-Annotation Quality Prediction for Semi-Supervised Learning with Ensembles
von: Simon, Dror, et al.
Veröffentlicht: (2019)
von: Simon, Dror, et al.
Veröffentlicht: (2019)
LESV: Language Embedded Sparse Voxel Fusion for Open-Vocabulary 3D Scene Understanding
von: Wang, Fusang, et al.
Veröffentlicht: (2026)
von: Wang, Fusang, et al.
Veröffentlicht: (2026)
Pedestrian Detection in Low-Light Conditions: A Comprehensive Survey
von: Ghari, Bahareh, et al.
Veröffentlicht: (2024)
von: Ghari, Bahareh, et al.
Veröffentlicht: (2024)
Black-box Adversarial Attacks on Monocular Depth Estimation Using Evolutionary Multi-objective Optimization
von: Daimo, Renya, et al.
Veröffentlicht: (2020)
von: Daimo, Renya, et al.
Veröffentlicht: (2020)
A Survey of Spatial Memory Representations for Efficient Robot Navigation
von: Pangaliman, Ma. Madecheen S., et al.
Veröffentlicht: (2026)
von: Pangaliman, Ma. Madecheen S., et al.
Veröffentlicht: (2026)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
UrbanAlign: Post-hoc Semantic Calibration for VLM-Human Preference Alignment
von: Zhang, Yecheng, et al.
Veröffentlicht: (2026)
von: Zhang, Yecheng, et al.
Veröffentlicht: (2026)
An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images
von: Castrillón-Santana, Modesto, et al.
Veröffentlicht: (2025)
von: Castrillón-Santana, Modesto, et al.
Veröffentlicht: (2025)
Textual and Visual Guided Task Adaptation for Source-Free Cross-Domain Few-Shot Segmentation
von: Liu, Jianming, et al.
Veröffentlicht: (2025)
von: Liu, Jianming, et al.
Veröffentlicht: (2025)
YotoR-You Only Transform One Representation
von: Villa, José Ignacio Díaz, et al.
Veröffentlicht: (2024)
von: Villa, José Ignacio Díaz, et al.
Veröffentlicht: (2024)
ERNet: Efficient Non-Rigid Registration Network for Point Sequences
von: He, Guangzhao, et al.
Veröffentlicht: (2025)
von: He, Guangzhao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Survey on Dynamic Neural Networks: from Computer Vision to Multi-modal Sensor Fusion
von: Montello, Fabio, et al.
Veröffentlicht: (2025) -
3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding
von: Chen, Yiping, et al.
Veröffentlicht: (2026) -
U-Net-Like Spiking Neural Networks for Single Image Dehazing
von: Li, Huibin, et al.
Veröffentlicht: (2025) -
VisRL: Intention-Driven Visual Perception via Reinforced Reasoning
von: Chen, Zhangquan, et al.
Veröffentlicht: (2025) -
4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026)