Look Before You Fuse: 2D-Guided Cross-Modal Alignment for Robust 3D Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xiang, Hu, Zhangchi, Xu, Xiao, Kong, Bin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hyperbolic Distillation: Geometry-Guided Cross-Modal Transfer for Robust 3D Object Detection
by: Ning, Kanglin, et al.
Published: (2026)
by: Ning, Kanglin, et al.
Published: (2026)
Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations
by: Kim, Jeonghyeon, et al.
Published: (2025)
by: Kim, Jeonghyeon, et al.
Published: (2025)
Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning
by: Li, Yian, et al.
Published: (2024)
by: Li, Yian, et al.
Published: (2024)
DPG-CD: Depth-Prior-Guided Cross-Modal Joint 2D-3D Change Detection
by: Zhang, Luqi, et al.
Published: (2026)
by: Zhang, Luqi, et al.
Published: (2026)
DiffCrossGait: Trajectory-Level Alignment for 2D-3D Cross-Modal Gait Recognition via Latent Diffusion
by: Lu, Zhiyang, et al.
Published: (2026)
by: Lu, Zhiyang, et al.
Published: (2026)
2D_3D Feature Fusion via Cross-Modal Latent Synthesis and Attention Guided Restoration for Industrial Anomaly Detection
by: Ali, Usman, et al.
Published: (2025)
by: Ali, Usman, et al.
Published: (2025)
Cross-Modal Mapping and Dual-Branch Reconstruction for 2D-3D Multimodal Industrial Anomaly Detection
by: Daci, Radia, et al.
Published: (2026)
by: Daci, Radia, et al.
Published: (2026)
AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
by: Zhang, Junyang, et al.
Published: (2025)
by: Zhang, Junyang, et al.
Published: (2025)
MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
by: Wang, Jiaze, et al.
Published: (2024)
by: Wang, Jiaze, et al.
Published: (2024)
Fuse Before Transfer: Knowledge Fusion for Heterogeneous Distillation
by: Li, Guopeng, et al.
Published: (2024)
by: Li, Guopeng, et al.
Published: (2024)
Large Language Model Guided Progressive Feature Alignment for Multimodal UAV Object Detection
by: Wu, Wentao, et al.
Published: (2025)
by: Wu, Wentao, et al.
Published: (2025)
Align Where the Words Look: Cross-Attention-Guided Patch Alignment with Contrastive and Transport Regularization for Bengali Captioning
by: Anonto, Riad Ahmed, et al.
Published: (2025)
by: Anonto, Riad Ahmed, et al.
Published: (2025)
Look Inside for More: Internal Spatial Modality Perception for 3D Anomaly Detection
by: Liang, Hanzhe, et al.
Published: (2024)
by: Liang, Hanzhe, et al.
Published: (2024)
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
by: Fang, Xiang, et al.
Published: (2022)
by: Fang, Xiang, et al.
Published: (2022)
AW-MoE: All-Weather Mixture of Experts for Robust Multi-Modal 3D Object Detection
by: Lin, Hongwei, et al.
Published: (2026)
by: Lin, Hongwei, et al.
Published: (2026)
SCKD: Semi-Supervised Cross-Modality Knowledge Distillation for 4D Radar Object Detection
by: Xu, Ruoyu, et al.
Published: (2024)
by: Xu, Ruoyu, et al.
Published: (2024)
Cross-Modal Purification and Fusion for Small-Object RGB-D Transmission-Line Defect Detection
by: Cui, Jiaming, et al.
Published: (2026)
by: Cui, Jiaming, et al.
Published: (2026)
Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination
by: Zheng, Haojie, et al.
Published: (2024)
by: Zheng, Haojie, et al.
Published: (2024)
Seeing is Believing: Robust Vision-Guided Cross-Modal Prompt Learning under Label Noise
by: Geng, Zibin, et al.
Published: (2026)
by: Geng, Zibin, et al.
Published: (2026)
Looking Back and Forth: Cross-Image Attention Calibration and Attentive Preference Learning for Multi-Image Hallucination Mitigation
by: Yang, Xiaochen, et al.
Published: (2026)
by: Yang, Xiaochen, et al.
Published: (2026)
From 2D Alignment to 3D Plausibility: Unifying Heterogeneous 2D Priors and Penetration-Free Diffusion for Occlusion-Robust Two-Hand Reconstruction
by: Han, Gaoge, et al.
Published: (2025)
by: Han, Gaoge, et al.
Published: (2025)
Depth-Guided Self-Supervised Human Keypoint Detection via Cross-Modal Distillation
by: Anand, Aman, et al.
Published: (2024)
by: Anand, Aman, et al.
Published: (2024)
Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis
by: Yu, Yang, et al.
Published: (2026)
by: Yu, Yang, et al.
Published: (2026)
Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models
by: Yan, Hanqi, et al.
Published: (2025)
by: Yan, Hanqi, et al.
Published: (2025)
SonarSweep: Fusing Sonar and Vision for Robust 3D Reconstruction via Plane Sweeping
by: Chen, Lingpeng, et al.
Published: (2025)
by: Chen, Lingpeng, et al.
Published: (2025)
UniPre3D: Unified Pre-training of 3D Point Cloud Models with Cross-Modal Gaussian Splatting
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
Multimodal Robust Prompt Distillation for 3D Point Cloud Models
by: Gu, Xiang, et al.
Published: (2025)
by: Gu, Xiang, et al.
Published: (2025)
DAE-Fuse: An Adaptive Discriminative Autoencoder for Multi-Modality Image Fusion
by: Guo, Yuchen, et al.
Published: (2024)
by: Guo, Yuchen, et al.
Published: (2024)
Capturing Fine-Grained Alignments Improves 3D Affordance Detection
by: Tokumitsu, Junsei, et al.
Published: (2025)
by: Tokumitsu, Junsei, et al.
Published: (2025)
MultiCorrupt: A Multi-Modal Robustness Dataset and Benchmark of LiDAR-Camera Fusion for 3D Object Detection
by: Beemelmanns, Till, et al.
Published: (2024)
by: Beemelmanns, Till, et al.
Published: (2024)
Robust Prior-Guided Segmentation for Editable 3D Gaussian Splatting
by: Joshi, Raushan, et al.
Published: (2026)
by: Joshi, Raushan, et al.
Published: (2026)
D3S2: Diffusion-Guided Dataset Distillation for Semantic Segmentation
by: Zheng, Wenjie, et al.
Published: (2026)
by: Zheng, Wenjie, et al.
Published: (2026)
Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language Models
by: Yu, Lu, et al.
Published: (2024)
by: Yu, Lu, et al.
Published: (2024)
CORENet: Cross-Modal 4D Radar Denoising Network with LiDAR Supervision for Autonomous Driving
by: Liu, Fuyang, et al.
Published: (2025)
by: Liu, Fuyang, et al.
Published: (2025)
What are You Looking at? Modality Contribution in Multimodal Medical Deep Learning
by: Gapp, Christian, et al.
Published: (2025)
by: Gapp, Christian, et al.
Published: (2025)
On the Adversarial Robustness of Camera-based 3D Object Detection
by: Xie, Shaoyuan, et al.
Published: (2023)
by: Xie, Shaoyuan, et al.
Published: (2023)
A 3D Generation Framework from Cross Modality to Parameterized Primitive
by: Liang, Yiming, et al.
Published: (2025)
by: Liang, Yiming, et al.
Published: (2025)
Diffusion-Based Restoration for Multi-Modal 3D Object Detection in Adverse Weather
by: He, Zhijian, et al.
Published: (2025)
by: He, Zhijian, et al.
Published: (2025)
S-LAM3D: Segmentation-Guided Monocular 3D Object Detection via Feature Space Fusion
by: Sas, Diana-Alexandra, et al.
Published: (2025)
by: Sas, Diana-Alexandra, et al.
Published: (2025)
Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles
by: Liao, Haicheng, et al.
Published: (2025)
by: Liao, Haicheng, et al.
Published: (2025)
Similar Items
-
Hyperbolic Distillation: Geometry-Guided Cross-Modal Transfer for Robust 3D Object Detection
by: Ning, Kanglin, et al.
Published: (2026) -
Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations
by: Kim, Jeonghyeon, et al.
Published: (2025) -
Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning
by: Li, Yian, et al.
Published: (2024) -
DPG-CD: Depth-Prior-Guided Cross-Modal Joint 2D-3D Change Detection
by: Zhang, Luqi, et al.
Published: (2026) -
DiffCrossGait: Trajectory-Level Alignment for 2D-3D Cross-Modal Gait Recognition via Latent Diffusion
by: Lu, Zhiyang, et al.
Published: (2026)