SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yuxuan, Li, Xiang, Li, Yunheng, Zhang, Yicheng, Dai, Yimian, Hou, Qibin, Cheng, Ming-Ming, Yang, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining
by: Li, Yuxuan, et al.
Published: (2026)
by: Li, Yuxuan, et al.
Published: (2026)
LSKNet: A Foundation Lightweight Backbone for Remote Sensing
by: Li, Yuxuan, et al.
Published: (2024)
by: Li, Yuxuan, et al.
Published: (2024)
DenoDet V2: Phase-Amplitude Cross Denoising for SAR Object Detection
by: Ni, Kang, et al.
Published: (2025)
by: Ni, Kang, et al.
Published: (2025)
Strip R-CNN: Large Strip Convolution for Remote Sensing Object Detection
by: Yuan, Xinbin, et al.
Published: (2025)
by: Yuan, Xinbin, et al.
Published: (2025)
DenoDet: Attention as Deformable Multi-Subspace Feature Denoising for Target Detection in SAR Images
by: Dai, Yimian, et al.
Published: (2024)
by: Dai, Yimian, et al.
Published: (2024)
HazyDet: Open-Source Benchmark for Drone-View Object Detection with Depth-Cues in Hazy Scenes
by: Feng, Changfeng, et al.
Published: (2024)
by: Feng, Changfeng, et al.
Published: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
by: Zhang, Zhenxing, et al.
Published: (2024)
by: Zhang, Zhenxing, et al.
Published: (2024)
Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing Images
by: Yuan, Bo, et al.
Published: (2024)
by: Yuan, Bo, et al.
Published: (2024)
Transcending Fusion: A Multi-Scale Alignment Method for Remote Sensing Image-Text Retrieval
by: Yang, Rui, et al.
Published: (2024)
by: Yang, Rui, et al.
Published: (2024)
Visual Instruction Pretraining for Domain-Specific Foundation Models
by: Li, Yuxuan, et al.
Published: (2025)
by: Li, Yuxuan, et al.
Published: (2025)
KAN-SAM: Kolmogorov-Arnold Network Guided Segment Anything Model for RGB-T Salient Object Detection
by: Li, Xingyuan, et al.
Published: (2025)
by: Li, Xingyuan, et al.
Published: (2025)
SatFusion: A Unified Framework for Enhancing Remote Sensing Images via Multi-Frame and Multi-Source Images Fusion
by: Tong, Yufei, et al.
Published: (2025)
by: Tong, Yufei, et al.
Published: (2025)
Copy-Move Forgery Detection and Question Answering for Remote Sensing Image
by: Zhang, Ze, et al.
Published: (2024)
by: Zhang, Ze, et al.
Published: (2024)
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
by: Miao, Bingchen, et al.
Published: (2024)
by: Miao, Bingchen, et al.
Published: (2024)
A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs
by: Dang, Yunkai, et al.
Published: (2025)
by: Dang, Yunkai, et al.
Published: (2025)
MambaPro: Multi-Modal Object Re-Identification with Mamba Aggregation and Synergistic Prompt
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
ContextDet: Temporal Action Detection with Adaptive Context Aggregation
by: Wang, Ning, et al.
Published: (2024)
by: Wang, Ning, et al.
Published: (2024)
Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding
by: Huang, Dawei, et al.
Published: (2025)
by: Huang, Dawei, et al.
Published: (2025)
Towards Unified Representation of Multi-Modal Pre-training for 3D Understanding via Differentiable Rendering
by: Fei, Ben, et al.
Published: (2024)
by: Fei, Ben, et al.
Published: (2024)
Unbiased Region-Language Alignment for Open-Vocabulary Dense Prediction
by: Li, Yunheng, et al.
Published: (2024)
by: Li, Yunheng, et al.
Published: (2024)
SARDet-100K: Towards Open-Source Benchmark and ToolKit for Large-Scale SAR Object Detection
by: Li, Yuxuan, et al.
Published: (2024)
by: Li, Yuxuan, et al.
Published: (2024)
DualComp: End-to-End Learning of a Unified Dual-Modality Lossless Compressor
by: Zhao, Yan, et al.
Published: (2025)
by: Zhao, Yan, et al.
Published: (2025)
MSCT: Differential Cross-Modal Attention for Deepfake Detection
by: Wei, Fangda, et al.
Published: (2026)
by: Wei, Fangda, et al.
Published: (2026)
Modality-Aware Shot Relating and Comparing for Video Scene Detection
by: Tan, Jiawei, et al.
Published: (2024)
by: Tan, Jiawei, et al.
Published: (2024)
ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling
by: Yang, Jianxuan, et al.
Published: (2026)
by: Yang, Jianxuan, et al.
Published: (2026)
M3FAS: An Accurate and Robust MultiModal Mobile Face Anti-Spoofing System
by: Kong, Chenqi, et al.
Published: (2023)
by: Kong, Chenqi, et al.
Published: (2023)
WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object Detection
by: Zhu, Haodong, et al.
Published: (2025)
by: Zhu, Haodong, et al.
Published: (2025)
Signal: Selective Interaction and Global-local Alignment for Multi-Modal Object Re-Identification
by: Liu, Yangyang, et al.
Published: (2025)
by: Liu, Yangyang, et al.
Published: (2025)
ISMAF: Intrinsic-Social Modality Alignment and Fusion for Multimodal Rumor Detection
by: Yu, Zihao, et al.
Published: (2025)
by: Yu, Zihao, et al.
Published: (2025)
Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing
by: Chen, Yaru, et al.
Published: (2025)
by: Chen, Yaru, et al.
Published: (2025)
SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
by: Zheng, Junjie, et al.
Published: (2025)
by: Zheng, Junjie, et al.
Published: (2025)
Towards Open-Vocabulary Remote Sensing Image Semantic Segmentation
by: Ye, Chengyang, et al.
Published: (2024)
by: Ye, Chengyang, et al.
Published: (2024)
Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer
by: Wang, Haoxu, et al.
Published: (2024)
by: Wang, Haoxu, et al.
Published: (2024)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
by: Chen, Liyang, et al.
Published: (2025)
by: Chen, Liyang, et al.
Published: (2025)
MTFusion: Reconstructing Any 3D Object from Single Image Using Multi-word Textual Inversion
by: Liu, Yu, et al.
Published: (2024)
by: Liu, Yu, et al.
Published: (2024)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
by: Chow, Wei, et al.
Published: (2024)
by: Chow, Wei, et al.
Published: (2024)
On the Robustness of Human-Object Interaction Detection against Distribution Shift
by: Xie, Chi, et al.
Published: (2025)
by: Xie, Chi, et al.
Published: (2025)
MB-ORES: A Multi-Branch Object Reasoner for Visual Grounding in Remote Sensing
by: Radouane, Karim, et al.
Published: (2025)
by: Radouane, Karim, et al.
Published: (2025)
CLOSP: A Unified Semantic Space for SAR, MSI, and Text in Remote Sensing
by: Cambrin, Daniele Rege, et al.
Published: (2025)
by: Cambrin, Daniele Rege, et al.
Published: (2025)
Similar Items
-
Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining
by: Li, Yuxuan, et al.
Published: (2026) -
LSKNet: A Foundation Lightweight Backbone for Remote Sensing
by: Li, Yuxuan, et al.
Published: (2024) -
DenoDet V2: Phase-Amplitude Cross Denoising for SAR Object Detection
by: Ni, Kang, et al.
Published: (2025) -
Strip R-CNN: Large Strip Convolution for Remote Sensing Object Detection
by: Yuan, Xinbin, et al.
Published: (2025) -
DenoDet: Attention as Deformable Multi-Subspace Feature Denoising for Target Detection in SAR Images
by: Dai, Yimian, et al.
Published: (2024)