Selection, Not Fusion: Radar-Modulated State Space Models for Radar-Camera Depth Estimation
Fuente:
arXiv
Guardado en:
| Autores principales: | Hou, Zhangcheng, Ohtsuki, Tomoaki |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Neuromorphic Monocular Depth Estimation with Uncertainty Modeling
por: Bergkvist, Viktor, et al.
Publicado: (2026)
por: Bergkvist, Viktor, et al.
Publicado: (2026)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
por: Raoufi, Behnam, et al.
Publicado: (2025)
por: Raoufi, Behnam, et al.
Publicado: (2025)
Smooth regularization for efficient video recognition
por: Goldman, Gil, et al.
Publicado: (2025)
por: Goldman, Gil, et al.
Publicado: (2025)
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
por: Durrani, Hamza Ahmed, et al.
Publicado: (2026)
por: Durrani, Hamza Ahmed, et al.
Publicado: (2026)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
por: Qesaraku, Bjorna, et al.
Publicado: (2025)
por: Qesaraku, Bjorna, et al.
Publicado: (2025)
FeedbackSTS-Det: Sparse Frames-Based Spatio-Temporal Semantic Feedback Network for Moving Infrared Small Target Detection
por: Huang, Yian, et al.
Publicado: (2026)
por: Huang, Yian, et al.
Publicado: (2026)
SERA-H: Beyond Native Sentinel Spatial Limits for High-Resolution Canopy Height Mapping
por: Boudras, Thomas, et al.
Publicado: (2025)
por: Boudras, Thomas, et al.
Publicado: (2025)
MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation
por: Bartkowiak, Patryk, et al.
Publicado: (2026)
por: Bartkowiak, Patryk, et al.
Publicado: (2026)
Lightweight Prompt-Guided CLIP Adaptation for Monocular Depth Estimation
por: Manghotay, Reyhaneh Ahani, et al.
Publicado: (2026)
por: Manghotay, Reyhaneh Ahani, et al.
Publicado: (2026)
OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models
por: Koroglu, Mathis, et al.
Publicado: (2024)
por: Koroglu, Mathis, et al.
Publicado: (2024)
Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object Detection
por: Lin, Xiaojian, et al.
Publicado: (2025)
por: Lin, Xiaojian, et al.
Publicado: (2025)
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
por: Wu, Jason, et al.
Publicado: (2026)
por: Wu, Jason, et al.
Publicado: (2026)
Perceptual Flow Network for Visually Grounded Reasoning
por: Li, Yangfu, et al.
Publicado: (2026)
por: Li, Yangfu, et al.
Publicado: (2026)
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
por: Liu, Bingnan, et al.
Publicado: (2026)
por: Liu, Bingnan, et al.
Publicado: (2026)
Caption-Driven Explainability: Probing CNNs for Bias via CLIP
por: Koller, Patrick, et al.
Publicado: (2025)
por: Koller, Patrick, et al.
Publicado: (2025)
Optimal Transport-Guided Source-Free Adaptation for Face Anti-Spoofing
por: Li, Zhuowei, et al.
Publicado: (2025)
por: Li, Zhuowei, et al.
Publicado: (2025)
Hierarchical Image-Guided 3D Point Cloud Segmentation in Industrial Scenes via Multi-View Bayesian Fusion
por: Zhu, Yu, et al.
Publicado: (2025)
por: Zhu, Yu, et al.
Publicado: (2025)
UnCageNet: Tracking and Pose Estimation of Caged Animal
por: Dutta, Sayak, et al.
Publicado: (2025)
por: Dutta, Sayak, et al.
Publicado: (2025)
Implementing Adaptations for Vision AutoRegressive Model
por: Shaikh, Kaif, et al.
Publicado: (2025)
por: Shaikh, Kaif, et al.
Publicado: (2025)
Single-Shot Metric Depth from Focused Plenoptic Cameras
por: Lasheras-Hernandez, Blanca, et al.
Publicado: (2024)
por: Lasheras-Hernandez, Blanca, et al.
Publicado: (2024)
Learning 3D object-centric representation through prediction
por: Day, John, et al.
Publicado: (2024)
por: Day, John, et al.
Publicado: (2024)
Efficient Attention: Attention with Linear Complexities
por: Shen, Zhuoran, et al.
Publicado: (2018)
por: Shen, Zhuoran, et al.
Publicado: (2018)
SPMamba-YOLO: An Underwater Object Detection Network Based on Multi-Scale Feature Enhancement and Global Context Modeling
por: Liao, Guanghao, et al.
Publicado: (2026)
por: Liao, Guanghao, et al.
Publicado: (2026)
Object detection in adverse weather conditions for autonomous vehicles using Instruct Pix2Pix
por: Gurbindo, Unai, et al.
Publicado: (2025)
por: Gurbindo, Unai, et al.
Publicado: (2025)
VLM-VPI: A Vision-Language Reasoning Framework for Improving Automated Vehicle-Pedestrian Interactions
por: Pu, Qingwen, et al.
Publicado: (2026)
por: Pu, Qingwen, et al.
Publicado: (2026)
Exploring Surround-View Fisheye Camera 3D Object Detection
por: Li, Changcai, et al.
Publicado: (2025)
por: Li, Changcai, et al.
Publicado: (2025)
GLoT: A Novel Gated-Logarithmic Transformer for Efficient Sign Language Translation
por: Shahin, Nada, et al.
Publicado: (2025)
por: Shahin, Nada, et al.
Publicado: (2025)
How Can One Choose the Best CAM-Based Explainability Method for a CNN Model?
por: Costa, Daniel da Silva, et al.
Publicado: (2026)
por: Costa, Daniel da Silva, et al.
Publicado: (2026)
An Analysis of Layer-Freezing Strategies for Enhanced Transfer Learning in YOLO Architectures
por: Dobrzycki, Andrzej D., et al.
Publicado: (2025)
por: Dobrzycki, Andrzej D., et al.
Publicado: (2025)
In Context Learning with Vision Transformers: Case Study
por: Zhao, Antony, et al.
Publicado: (2025)
por: Zhao, Antony, et al.
Publicado: (2025)
IMKD: Intensity-Aware Multi-Level Knowledge Distillation for Camera-Radar Fusion
por: Mishra, Shashank, et al.
Publicado: (2025)
por: Mishra, Shashank, et al.
Publicado: (2025)
SpectralCA: Bi-Directional Cross-Attention for Next-Generation UAV Hyperspectral Vision
por: Brovko, D. V.
Publicado: (2025)
por: Brovko, D. V.
Publicado: (2025)
Do Generative Metrics Predict YOLO Performance? An Evaluation Across Models, Augmentation Ratios, and Dataset Complexity
por: Marian, Vasile, et al.
Publicado: (2026)
por: Marian, Vasile, et al.
Publicado: (2026)
Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D
por: Arnaud, Sergio, et al.
Publicado: (2025)
por: Arnaud, Sergio, et al.
Publicado: (2025)
Towards a Generalizable Fusion Architecture for Multimodal Object Detection
por: Berjawi, Jad, et al.
Publicado: (2025)
por: Berjawi, Jad, et al.
Publicado: (2025)
THIRDEYE: Cue-Aware Monocular Depth Estimation via Brain-Inspired Multi-Stage Fusion
por: Ioan, Calin Teodor
Publicado: (2025)
por: Ioan, Calin Teodor
Publicado: (2025)
ADAT: Time-Series-Aware Adaptive Transformer Architecture for Sign Language Translation
por: Shahin, Nada, et al.
Publicado: (2025)
por: Shahin, Nada, et al.
Publicado: (2025)
Spiking Neural Networks for event-based action recognition: A new task to understand their advantage
por: Vicente-Sola, Alex, et al.
Publicado: (2022)
por: Vicente-Sola, Alex, et al.
Publicado: (2022)
Deep Learning-based Depth Estimation Methods from Monocular Image and Videos: A Comprehensive Survey
por: Rajapaksha, Uchitha, et al.
Publicado: (2024)
por: Rajapaksha, Uchitha, et al.
Publicado: (2024)
Car Object Counting and Position Estimation via Extension of the CLIP-EBC Framework
por: Jung, Seoik, et al.
Publicado: (2025)
por: Jung, Seoik, et al.
Publicado: (2025)
Ejemplares similares
-
Neuromorphic Monocular Depth Estimation with Uncertainty Modeling
por: Bergkvist, Viktor, et al.
Publicado: (2026) -
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
por: Raoufi, Behnam, et al.
Publicado: (2025) -
Smooth regularization for efficient video recognition
por: Goldman, Gil, et al.
Publicado: (2025) -
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
por: Durrani, Hamza Ahmed, et al.
Publicado: (2026) -
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
por: Qesaraku, Bjorna, et al.
Publicado: (2025)