Rethinking Multi-Modal Object Detection from the Perspective of Mono-Modality Feature Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhao, Tianyi, Liu, Boyang, Gao, Yanglei, Sun, Yiming, Yuan, Maoxun, Wei, Xingxing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Removal then Selection: A Coarse-to-Fine Fusion Perspective for RGB-Infrared Object Detection
por: Zhao, Tianyi, et al.
Publicado: (2024)
por: Zhao, Tianyi, et al.
Publicado: (2024)
$\mathbf{C}^2$Former: Calibrated and Complementary Transformer for RGB-Infrared Object Detection
por: Yuan, Maoxun, et al.
Publicado: (2023)
por: Yuan, Maoxun, et al.
Publicado: (2023)
Breaking Self-Attention Failure: Rethinking Query Initialization for Infrared Small Target Detection
por: Liu, Yuteng, et al.
Publicado: (2026)
por: Liu, Yuteng, et al.
Publicado: (2026)
Seeing Through the Noise: Improving Infrared Small Target Detection and Segmentation from Noise Suppression Perspective
por: Yuan, Maoxun, et al.
Publicado: (2025)
por: Yuan, Maoxun, et al.
Publicado: (2025)
UniRGB-IR: A Unified Framework for Visible-Infrared Semantic Tasks via Adapter Tuning
por: Yuan, Maoxun, et al.
Publicado: (2024)
por: Yuan, Maoxun, et al.
Publicado: (2024)
RGBT-Ground Benchmark: Visual Grounding Beyond RGB in Complex Real-World Scenarios
por: Zhao, Tianyi, et al.
Publicado: (2025)
por: Zhao, Tianyi, et al.
Publicado: (2025)
Exploiting Modality-Specific Features For Multi-Modal Manipulation Detection And Grounding
por: Wang, Jiazhen, et al.
Publicado: (2023)
por: Wang, Jiazhen, et al.
Publicado: (2023)
Modality-Agnostic Prompt Learning for Multi-Modal Camouflaged Object Detection
por: Wang, Hao, et al.
Publicado: (2026)
por: Wang, Hao, et al.
Publicado: (2026)
DEYOLO: Dual-Feature-Enhancement YOLO for Cross-Modality Object Detection
por: Chen, Yishuo, et al.
Publicado: (2024)
por: Chen, Yishuo, et al.
Publicado: (2024)
MV2DFusion: Leveraging Modality-Specific Object Semantics for Multi-Modal 3D Detection
por: Wang, Zitian, et al.
Publicado: (2024)
por: Wang, Zitian, et al.
Publicado: (2024)
Knowledge-Guided Adversarial Training for Infrared Object Detection via Thermal Radiation Modeling
por: Zhao, Shiji, et al.
Publicado: (2026)
por: Zhao, Shiji, et al.
Publicado: (2026)
M^3Detection: Multi-Frame Multi-Level Feature Fusion for Multi-Modal 3D Object Detection with Camera and 4D Imaging Radar
por: Li, Xiaozhi, et al.
Publicado: (2025)
por: Li, Xiaozhi, et al.
Publicado: (2025)
Modality Prompts for Arbitrary Modality Salient Object Detection
por: Huang, Nianchang, et al.
Publicado: (2024)
por: Huang, Nianchang, et al.
Publicado: (2024)
UniMMAD: Unified Multi-Modal and Multi-Class Anomaly Detection via MoE-Driven Feature Decompression
por: Zhao, Yuan, et al.
Publicado: (2025)
por: Zhao, Yuan, et al.
Publicado: (2025)
GraphAlign: Enhancing Accurate Feature Alignment by Graph matching for Multi-Modal 3D Object Detection
por: Song, Ziying, et al.
Publicado: (2023)
por: Song, Ziying, et al.
Publicado: (2023)
GraphBEV: Towards Robust BEV Feature Alignment for Multi-Modal 3D Object Detection
por: Song, Ziying, et al.
Publicado: (2024)
por: Song, Ziying, et al.
Publicado: (2024)
COMMA: Co-Articulated Multi-Modal Learning
por: Hu, Lianyu, et al.
Publicado: (2023)
por: Hu, Lianyu, et al.
Publicado: (2023)
Exploring Modality-Aware Fusion and Decoupled Temporal Propagation for Multi-Modal Object Tracking
por: Wang, Shilei, et al.
Publicado: (2026)
por: Wang, Shilei, et al.
Publicado: (2026)
ContrastAlign: Toward Robust BEV Feature Alignment via Contrastive Learning for Multi-Modal 3D Object Detection
por: Song, Ziying, et al.
Publicado: (2024)
por: Song, Ziying, et al.
Publicado: (2024)
Learning Multi-Modal Prototypes for Cross-Domain Few-Shot Object Detection
por: Wang, Wanqi, et al.
Publicado: (2026)
por: Wang, Wanqi, et al.
Publicado: (2026)
ModalPatch: A Plug-and-Play Module for Robust Multi-Modal 3D Object Detection under Modality Drop
por: Li, Shuangzhi, et al.
Publicado: (2026)
por: Li, Shuangzhi, et al.
Publicado: (2026)
Multi-Modal Guided Multi-Source Domain Adaptation for Object Detection
por: Lee, Sangin, et al.
Publicado: (2026)
por: Lee, Sangin, et al.
Publicado: (2026)
Integrating Object Detection Modality into Visual Language Model for Enhanced Autonomous Driving Agent
por: He, Linfeng, et al.
Publicado: (2024)
por: He, Linfeng, et al.
Publicado: (2024)
CalFuse: Multi-Modal Continual Learning via Feature Calibration and Parameter Fusion
por: Guo, Juncen, et al.
Publicado: (2025)
por: Guo, Juncen, et al.
Publicado: (2025)
CMF-IoU: Multi-Stage Cross-Modal Fusion 3D Object Detection with IoU Joint Prediction
por: Ning, Zhiwei, et al.
Publicado: (2025)
por: Ning, Zhiwei, et al.
Publicado: (2025)
MDReID: Modality-Decoupled Learning for Any-to-Any Multi-Modal Object Re-Identification
por: Feng, Yingying, et al.
Publicado: (2025)
por: Feng, Yingying, et al.
Publicado: (2025)
STMI: Segmentation-Guided Token Modulation with Cross-Modal Hypergraph Interaction for Multi-Modal Object Re-Identification
por: Xu, Xingguo, et al.
Publicado: (2026)
por: Xu, Xingguo, et al.
Publicado: (2026)
MIFNet: Learning Modality-Invariant Features for Generalizable Multimodal Image Matching
por: Liu, Yepeng, et al.
Publicado: (2025)
por: Liu, Yepeng, et al.
Publicado: (2025)
DeMo: Decoupled Feature-Based Mixture of Experts for Multi-Modal Object Re-Identification
por: Wang, Yuhao, et al.
Publicado: (2024)
por: Wang, Yuhao, et al.
Publicado: (2024)
DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection
por: Gare, Gautam Rajendrakumar, et al.
Publicado: (2026)
por: Gare, Gautam Rajendrakumar, et al.
Publicado: (2026)
Multi-Modal Assistance for Unsupervised Domain Adaptation on Point Cloud 3D Object Detection
por: Zhao, Shenao, et al.
Publicado: (2025)
por: Zhao, Shenao, et al.
Publicado: (2025)
Progressive Multi-Modal Fusion for Robust 3D Object Detection
por: Mohan, Rohit, et al.
Publicado: (2024)
por: Mohan, Rohit, et al.
Publicado: (2024)
DGFusion: Dual-guided Fusion for Robust Multi-Modal 3D Object Detection
por: Jia, Feiyang, et al.
Publicado: (2025)
por: Jia, Feiyang, et al.
Publicado: (2025)
M2I2HA: Multi-modal Object Detection Based on Intra- and Inter-Modal Hypergraph Attention
por: Yang, Xiaofan, et al.
Publicado: (2026)
por: Yang, Xiaofan, et al.
Publicado: (2026)
DM$^3$T: Harmonizing Modalities via Diffusion for Multi-Object Tracking
por: Li, Weiran, et al.
Publicado: (2025)
por: Li, Weiran, et al.
Publicado: (2025)
Object-X: Learning to Reconstruct Multi-Modal 3D Object Representations
por: Di Lorenzo, Gaia, et al.
Publicado: (2025)
por: Di Lorenzo, Gaia, et al.
Publicado: (2025)
CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection
por: Wu, Yuchen, et al.
Publicado: (2026)
por: Wu, Yuchen, et al.
Publicado: (2026)
TUNI: Real-time RGB-T Semantic Segmentation with Unified Multi-Modal Feature Extraction and Cross-Modal Feature Fusion
por: Guo, Xiaodong, et al.
Publicado: (2025)
por: Guo, Xiaodong, et al.
Publicado: (2025)
Modality-Aware Feature Matching: A Comprehensive Review of Single- and Cross-Modality Techniques
por: Liu, Weide, et al.
Publicado: (2025)
por: Liu, Weide, et al.
Publicado: (2025)
Multi-Modal Face Anti-Spoofing via Cross-Modal Feature Transitions
por: Chong, Jun-Xiong, et al.
Publicado: (2025)
por: Chong, Jun-Xiong, et al.
Publicado: (2025)
Ejemplares similares
-
Removal then Selection: A Coarse-to-Fine Fusion Perspective for RGB-Infrared Object Detection
por: Zhao, Tianyi, et al.
Publicado: (2024) -
$\mathbf{C}^2$Former: Calibrated and Complementary Transformer for RGB-Infrared Object Detection
por: Yuan, Maoxun, et al.
Publicado: (2023) -
Breaking Self-Attention Failure: Rethinking Query Initialization for Infrared Small Target Detection
por: Liu, Yuteng, et al.
Publicado: (2026) -
Seeing Through the Noise: Improving Infrared Small Target Detection and Segmentation from Noise Suppression Perspective
por: Yuan, Maoxun, et al.
Publicado: (2025) -
UniRGB-IR: A Unified Framework for Visible-Infrared Semantic Tasks via Adapter Tuning
por: Yuan, Maoxun, et al.
Publicado: (2024)