KAN-RCBEVDepth: A multi-modal fusion algorithm in object detection for autonomous driving
Fuente:
arXiv
Guardado en:
| Autores principales: | Lai, Zhihao, Liu, Chuanhao, Sheng, Shihui, Zhang, Zhiqiang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Msmsfnet: a multi-stream and multi-scale fusion net for edge detection
por: Liu, Chenguang, et al.
Publicado: (2024)
por: Liu, Chenguang, et al.
Publicado: (2024)
Timealign: A multi-modal object detection method for time misalignment fusing in autonomous driving
por: Song, Zhihang, et al.
Publicado: (2024)
por: Song, Zhihang, et al.
Publicado: (2024)
A re-calibration method for object detection with multi-modal alignment bias in autonomous driving
por: Song, Zhihang, et al.
Publicado: (2024)
por: Song, Zhihang, et al.
Publicado: (2024)
Multi-model approach for autonomous driving: A comprehensive study on traffic sign-, vehicle- and lane detection and behavioral cloning
por: Jaisankar, Kanishkha, et al.
Publicado: (2026)
por: Jaisankar, Kanishkha, et al.
Publicado: (2026)
SToRM: Supervised Token Reduction for Multi-modal LLMs toward efficient end-to-end autonomous driving
por: Kim, Seo Hyun, et al.
Publicado: (2026)
por: Kim, Seo Hyun, et al.
Publicado: (2026)
PMMD: A pose-guided multi-view multi-modal diffusion for person generation
por: Shang, Ziyu, et al.
Publicado: (2025)
por: Shang, Ziyu, et al.
Publicado: (2025)
Demystifying KAN for Vision Tasks: The RepKAN Approach
por: Cheon, Minjong
Publicado: (2026)
por: Cheon, Minjong
Publicado: (2026)
KAN See In the Dark
por: Ning, Aoxiang, et al.
Publicado: (2024)
por: Ning, Aoxiang, et al.
Publicado: (2024)
Physics-Consistent Diffusion for Efficient Fluid Super-Resolution via Multiscale Residual Correction
por: Li, Zhihao, et al.
Publicado: (2026)
por: Li, Zhihao, et al.
Publicado: (2026)
DM-QPMNET: Dual-modality fusion network for cell segmentation in quantitative phase microscopy
por: Chakraborty, Rajatsubhra, et al.
Publicado: (2025)
por: Chakraborty, Rajatsubhra, et al.
Publicado: (2025)
EyeCLIP: A visual-language foundation model for multi-modal ophthalmic image analysis
por: Shi, Danli, et al.
Publicado: (2024)
por: Shi, Danli, et al.
Publicado: (2024)
R2Det: Exploring Relaxed Rotation Equivariance in 2D object detection
por: Wu, Zhiqiang, et al.
Publicado: (2024)
por: Wu, Zhiqiang, et al.
Publicado: (2024)
K-U-KAN: Koopman-Enhanced U-KAN for 3D Dental Reconstruction from a Single Panoramic X-ray Radiograph
por: Parida, Bikram Keshari, et al.
Publicado: (2026)
por: Parida, Bikram Keshari, et al.
Publicado: (2026)
UAV traffic scene understanding: A regulation embedded multi-modal network and a unified benchmark
por: Zhang, Yu, et al.
Publicado: (2026)
por: Zhang, Yu, et al.
Publicado: (2026)
Multi-modal user interface control detection using cross-attention
por: Moradi, Milad, et al.
Publicado: (2026)
por: Moradi, Milad, et al.
Publicado: (2026)
Multi-Branch Auxiliary Fusion YOLO with Re-parameterization Heterogeneous Convolutional for accurate object detection
por: Yang, Zhiqiang, et al.
Publicado: (2024)
por: Yang, Zhiqiang, et al.
Publicado: (2024)
Fine-grained Action Analysis: A Multi-modality and Multi-task Dataset of Figure Skating
por: Liu, Sheng-Lan, et al.
Publicado: (2023)
por: Liu, Sheng-Lan, et al.
Publicado: (2023)
Adaptive H&E-IHC information fusion staining framework based on feature extra
por: Jia, Yifan, et al.
Publicado: (2025)
por: Jia, Yifan, et al.
Publicado: (2025)
TransMA: an explainable multi-modal deep learning model for predicting properties of ionizable lipid nanoparticles in mRNA delivery
por: Wu, Kun, et al.
Publicado: (2024)
por: Wu, Kun, et al.
Publicado: (2024)
KAN-Mixers: a new deep learning architecture for image classification
por: Canuto, Jorge Luiz dos Santos, et al.
Publicado: (2025)
por: Canuto, Jorge Luiz dos Santos, et al.
Publicado: (2025)
KAN You See It? KANs and Sentinel for Effective and Explainable Crop Field Segmentation
por: Cambrin, Daniele Rege, et al.
Publicado: (2024)
por: Cambrin, Daniele Rege, et al.
Publicado: (2024)
Awesome Multi-modal Object Tracking
por: Zhang, Chunhui, et al.
Publicado: (2024)
por: Zhang, Chunhui, et al.
Publicado: (2024)
IA-T2I: Internet-Augmented Text-to-Image Generation
por: Li, Chuanhao, et al.
Publicado: (2025)
por: Li, Chuanhao, et al.
Publicado: (2025)
Integrating Text and Image Pre-training for Multi-modal Algorithmic Reasoning
por: Zhang, Zijian, et al.
Publicado: (2024)
por: Zhang, Zijian, et al.
Publicado: (2024)
Hierarchical Multi-modal Transformer for Cross-modal Long Document Classification
por: Liu, Tengfei, et al.
Publicado: (2024)
por: Liu, Tengfei, et al.
Publicado: (2024)
PatchDenoiser: Parameter-efficient multi-scale patch learning and fusion denoiser for Low-dose CT imaging
por: Fartiyal, Jitindra, et al.
Publicado: (2026)
por: Fartiyal, Jitindra, et al.
Publicado: (2026)
BFA-YOLO: A balanced multiscale object detection network for building façade attachments detection
por: Chen, Yangguang, et al.
Publicado: (2024)
por: Chen, Yangguang, et al.
Publicado: (2024)
Multi-Sourced Compositional Generalization in Visual Question Answering
por: Li, Chuanhao, et al.
Publicado: (2025)
por: Li, Chuanhao, et al.
Publicado: (2025)
SEATrack: Simple, Efficient, and Adaptive Multimodal Tracker
por: Su, Junbin, et al.
Publicado: (2026)
por: Su, Junbin, et al.
Publicado: (2026)
From classical techniques to convolution-based models: A review of object detection algorithms
por: Neha, Fnu, et al.
Publicado: (2024)
por: Neha, Fnu, et al.
Publicado: (2024)
Fusion-Mamba for Cross-modality Object Detection
por: Dong, Wenhao, et al.
Publicado: (2024)
por: Dong, Wenhao, et al.
Publicado: (2024)
SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge
por: Li, Chuanhao, et al.
Publicado: (2024)
por: Li, Chuanhao, et al.
Publicado: (2024)
Research on target detection method of distracted driving behavior based on improved YOLOv8
por: Shen, Shiquan, et al.
Publicado: (2024)
por: Shen, Shiquan, et al.
Publicado: (2024)
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
por: Liu, Chonghan, et al.
Publicado: (2025)
por: Liu, Chonghan, et al.
Publicado: (2025)
UNetVL: Enhancing 3D Medical Image Segmentation with Chebyshev KAN Powered Vision-LSTM
por: Guo, Xuhui, et al.
Publicado: (2025)
por: Guo, Xuhui, et al.
Publicado: (2025)
PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference
por: Mao, Xiaofeng, et al.
Publicado: (2026)
por: Mao, Xiaofeng, et al.
Publicado: (2026)
SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model
por: Chang, Yifan, et al.
Publicado: (2025)
por: Chang, Yifan, et al.
Publicado: (2025)
The detection and rectification for identity-switch based on unfalsified control
por: Huang, Junchao, et al.
Publicado: (2023)
por: Huang, Junchao, et al.
Publicado: (2023)
VLA-Mark: A cross modal watermark for large vision-language alignment model
por: Liu, Shuliang, et al.
Publicado: (2025)
por: Liu, Shuliang, et al.
Publicado: (2025)
A Novel Approach to for Multimodal Emotion Recognition : Multimodal semantic information fusion
por: Dai, Wei, et al.
Publicado: (2025)
por: Dai, Wei, et al.
Publicado: (2025)
Ejemplares similares
-
Msmsfnet: a multi-stream and multi-scale fusion net for edge detection
por: Liu, Chenguang, et al.
Publicado: (2024) -
Timealign: A multi-modal object detection method for time misalignment fusing in autonomous driving
por: Song, Zhihang, et al.
Publicado: (2024) -
A re-calibration method for object detection with multi-modal alignment bias in autonomous driving
por: Song, Zhihang, et al.
Publicado: (2024) -
Multi-model approach for autonomous driving: A comprehensive study on traffic sign-, vehicle- and lane detection and behavioral cloning
por: Jaisankar, Kanishkha, et al.
Publicado: (2026) -
SToRM: Supervised Token Reduction for Multi-modal LLMs toward efficient end-to-end autonomous driving
por: Kim, Seo Hyun, et al.
Publicado: (2026)