KAN-RCBEVDepth: A multi-modal fusion algorithm in object detection for autonomous driving
Fuente:
arXiv
Saved in:
| Main Authors: | Lai, Zhihao, Liu, Chuanhao, Sheng, Shihui, Zhang, Zhiqiang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Msmsfnet: a multi-stream and multi-scale fusion net for edge detection
by: Liu, Chenguang, et al.
Published: (2024)
by: Liu, Chenguang, et al.
Published: (2024)
Timealign: A multi-modal object detection method for time misalignment fusing in autonomous driving
by: Song, Zhihang, et al.
Published: (2024)
by: Song, Zhihang, et al.
Published: (2024)
A re-calibration method for object detection with multi-modal alignment bias in autonomous driving
by: Song, Zhihang, et al.
Published: (2024)
by: Song, Zhihang, et al.
Published: (2024)
Multi-model approach for autonomous driving: A comprehensive study on traffic sign-, vehicle- and lane detection and behavioral cloning
by: Jaisankar, Kanishkha, et al.
Published: (2026)
by: Jaisankar, Kanishkha, et al.
Published: (2026)
SToRM: Supervised Token Reduction for Multi-modal LLMs toward efficient end-to-end autonomous driving
by: Kim, Seo Hyun, et al.
Published: (2026)
by: Kim, Seo Hyun, et al.
Published: (2026)
Demystifying KAN for Vision Tasks: The RepKAN Approach
by: Cheon, Minjong
Published: (2026)
by: Cheon, Minjong
Published: (2026)
PMMD: A pose-guided multi-view multi-modal diffusion for person generation
by: Shang, Ziyu, et al.
Published: (2025)
by: Shang, Ziyu, et al.
Published: (2025)
KAN See In the Dark
by: Ning, Aoxiang, et al.
Published: (2024)
by: Ning, Aoxiang, et al.
Published: (2024)
Physics-Consistent Diffusion for Efficient Fluid Super-Resolution via Multiscale Residual Correction
by: Li, Zhihao, et al.
Published: (2026)
by: Li, Zhihao, et al.
Published: (2026)
DM-QPMNET: Dual-modality fusion network for cell segmentation in quantitative phase microscopy
by: Chakraborty, Rajatsubhra, et al.
Published: (2025)
by: Chakraborty, Rajatsubhra, et al.
Published: (2025)
EyeCLIP: A visual-language foundation model for multi-modal ophthalmic image analysis
by: Shi, Danli, et al.
Published: (2024)
by: Shi, Danli, et al.
Published: (2024)
R2Det: Exploring Relaxed Rotation Equivariance in 2D object detection
by: Wu, Zhiqiang, et al.
Published: (2024)
by: Wu, Zhiqiang, et al.
Published: (2024)
K-U-KAN: Koopman-Enhanced U-KAN for 3D Dental Reconstruction from a Single Panoramic X-ray Radiograph
by: Parida, Bikram Keshari, et al.
Published: (2026)
by: Parida, Bikram Keshari, et al.
Published: (2026)
UAV traffic scene understanding: A regulation embedded multi-modal network and a unified benchmark
by: Zhang, Yu, et al.
Published: (2026)
by: Zhang, Yu, et al.
Published: (2026)
Multi-modal user interface control detection using cross-attention
by: Moradi, Milad, et al.
Published: (2026)
by: Moradi, Milad, et al.
Published: (2026)
Multi-Branch Auxiliary Fusion YOLO with Re-parameterization Heterogeneous Convolutional for accurate object detection
by: Yang, Zhiqiang, et al.
Published: (2024)
by: Yang, Zhiqiang, et al.
Published: (2024)
Fine-grained Action Analysis: A Multi-modality and Multi-task Dataset of Figure Skating
by: Liu, Sheng-Lan, et al.
Published: (2023)
by: Liu, Sheng-Lan, et al.
Published: (2023)
Adaptive H&E-IHC information fusion staining framework based on feature extra
by: Jia, Yifan, et al.
Published: (2025)
by: Jia, Yifan, et al.
Published: (2025)
TransMA: an explainable multi-modal deep learning model for predicting properties of ionizable lipid nanoparticles in mRNA delivery
by: Wu, Kun, et al.
Published: (2024)
by: Wu, Kun, et al.
Published: (2024)
KAN-Mixers: a new deep learning architecture for image classification
by: Canuto, Jorge Luiz dos Santos, et al.
Published: (2025)
by: Canuto, Jorge Luiz dos Santos, et al.
Published: (2025)
KAN You See It? KANs and Sentinel for Effective and Explainable Crop Field Segmentation
by: Cambrin, Daniele Rege, et al.
Published: (2024)
by: Cambrin, Daniele Rege, et al.
Published: (2024)
Awesome Multi-modal Object Tracking
by: Zhang, Chunhui, et al.
Published: (2024)
by: Zhang, Chunhui, et al.
Published: (2024)
IA-T2I: Internet-Augmented Text-to-Image Generation
by: Li, Chuanhao, et al.
Published: (2025)
by: Li, Chuanhao, et al.
Published: (2025)
Hierarchical Multi-modal Transformer for Cross-modal Long Document Classification
by: Liu, Tengfei, et al.
Published: (2024)
by: Liu, Tengfei, et al.
Published: (2024)
Integrating Text and Image Pre-training for Multi-modal Algorithmic Reasoning
by: Zhang, Zijian, et al.
Published: (2024)
by: Zhang, Zijian, et al.
Published: (2024)
PatchDenoiser: Parameter-efficient multi-scale patch learning and fusion denoiser for Low-dose CT imaging
by: Fartiyal, Jitindra, et al.
Published: (2026)
by: Fartiyal, Jitindra, et al.
Published: (2026)
Multi-Sourced Compositional Generalization in Visual Question Answering
by: Li, Chuanhao, et al.
Published: (2025)
by: Li, Chuanhao, et al.
Published: (2025)
BFA-YOLO: A balanced multiscale object detection network for building façade attachments detection
by: Chen, Yangguang, et al.
Published: (2024)
by: Chen, Yangguang, et al.
Published: (2024)
SEATrack: Simple, Efficient, and Adaptive Multimodal Tracker
by: Su, Junbin, et al.
Published: (2026)
by: Su, Junbin, et al.
Published: (2026)
From classical techniques to convolution-based models: A review of object detection algorithms
by: Neha, Fnu, et al.
Published: (2024)
by: Neha, Fnu, et al.
Published: (2024)
Fusion-Mamba for Cross-modality Object Detection
by: Dong, Wenhao, et al.
Published: (2024)
by: Dong, Wenhao, et al.
Published: (2024)
SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge
by: Li, Chuanhao, et al.
Published: (2024)
by: Li, Chuanhao, et al.
Published: (2024)
Research on target detection method of distracted driving behavior based on improved YOLOv8
by: Shen, Shiquan, et al.
Published: (2024)
by: Shen, Shiquan, et al.
Published: (2024)
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
by: Liu, Chonghan, et al.
Published: (2025)
by: Liu, Chonghan, et al.
Published: (2025)
UNetVL: Enhancing 3D Medical Image Segmentation with Chebyshev KAN Powered Vision-LSTM
by: Guo, Xuhui, et al.
Published: (2025)
by: Guo, Xuhui, et al.
Published: (2025)
PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference
by: Mao, Xiaofeng, et al.
Published: (2026)
by: Mao, Xiaofeng, et al.
Published: (2026)
SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model
by: Chang, Yifan, et al.
Published: (2025)
by: Chang, Yifan, et al.
Published: (2025)
The detection and rectification for identity-switch based on unfalsified control
by: Huang, Junchao, et al.
Published: (2023)
by: Huang, Junchao, et al.
Published: (2023)
VLA-Mark: A cross modal watermark for large vision-language alignment model
by: Liu, Shuliang, et al.
Published: (2025)
by: Liu, Shuliang, et al.
Published: (2025)
TrajFlow: Multi-modal Motion Prediction via Flow Matching
by: Yan, Qi, et al.
Published: (2025)
by: Yan, Qi, et al.
Published: (2025)
Similar Items
-
Msmsfnet: a multi-stream and multi-scale fusion net for edge detection
by: Liu, Chenguang, et al.
Published: (2024) -
Timealign: A multi-modal object detection method for time misalignment fusing in autonomous driving
by: Song, Zhihang, et al.
Published: (2024) -
A re-calibration method for object detection with multi-modal alignment bias in autonomous driving
by: Song, Zhihang, et al.
Published: (2024) -
Multi-model approach for autonomous driving: A comprehensive study on traffic sign-, vehicle- and lane detection and behavioral cloning
by: Jaisankar, Kanishkha, et al.
Published: (2026) -
SToRM: Supervised Token Reduction for Multi-modal LLMs toward efficient end-to-end autonomous driving
by: Kim, Seo Hyun, et al.
Published: (2026)