E2E-MFD: Towards End-to-End Synchronous Multimodal Fusion Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jiaqing, Cao, Mingxiang, Xie, Weiying, Lei, Jie, Li, Daixun, Huang, Wenbo, Li, Yunsong, Yang, Xue |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-scale direction-aware SAR object detection network via global information fusion
by: Cao, Mingxiang, et al.
Published: (2023)
by: Cao, Mingxiang, et al.
Published: (2023)
Multimodal Informative ViT: Information Aggregation and Distribution for Hyperspectral and LiDAR Classification
by: Zhang, Jiaqing, et al.
Published: (2024)
by: Zhang, Jiaqing, et al.
Published: (2024)
FusionSAM: Visual Multi-Modal Learning with Segment Anything
by: Li, Daixun, et al.
Published: (2024)
by: Li, Daixun, et al.
Published: (2024)
Distribution-aware Interactive Attention Network and Large-scale Cloud Recognition Benchmark on FY-4A Satellite Image
by: Zhang, Jiaqing, et al.
Published: (2024)
by: Zhang, Jiaqing, et al.
Published: (2024)
M$^3$amba: CLIP-driven Mamba Model for Multi-modal Remote Sensing Classification
by: Cao, Mingxiang, et al.
Published: (2025)
by: Cao, Mingxiang, et al.
Published: (2025)
DiffCLIP: Few-shot Language-driven Multimodal Classifier
by: Zhang, Jiaqing, et al.
Published: (2024)
by: Zhang, Jiaqing, et al.
Published: (2024)
Reducing Spurious Correlation for Federated Domain Generalization
by: Ma, Shuran, et al.
Published: (2024)
by: Ma, Shuran, et al.
Published: (2024)
RS-DGC: Exploring Neighborhood Statistics for Dynamic Gradient Compression on Remote Sensing Image Interpretation
by: Xie, Weiying, et al.
Published: (2023)
by: Xie, Weiying, et al.
Published: (2023)
SeaDATE: Remedy Dual-Attention Transformer with Semantic Alignment via Contrast Learning for Multimodal Object Detection
by: Dong, Shuhan, et al.
Published: (2024)
by: Dong, Shuhan, et al.
Published: (2024)
SwiMDiff: Scene-wide Matching Contrastive Learning with Diffusion Constraint for Remote Sensing Image
by: Tian, Jiayuan, et al.
Published: (2024)
by: Tian, Jiayuan, et al.
Published: (2024)
FoRA: Low-Rank Adaptation Model beyond Multimodal Siamese Network
by: Xie, Weiying, et al.
Published: (2024)
by: Xie, Weiying, et al.
Published: (2024)
Exploring Hyperspectral Anomaly Detection with Human Vision: A Small Target Aware Detector
by: Ma, Jitao, et al.
Published: (2024)
by: Ma, Jitao, et al.
Published: (2024)
Hyperspectral Anomaly Detection with Self-Supervised Anomaly Prior
by: Liu, Yidan, et al.
Published: (2024)
by: Liu, Yidan, et al.
Published: (2024)
BSDM: Background Suppression Diffusion Model for Hyperspectral Anomaly Detection
by: Ma, Jitao, et al.
Published: (2023)
by: Ma, Jitao, et al.
Published: (2023)
Domain Adaptation for Large-Vocabulary Object Detectors
by: Jiang, Kai, et al.
Published: (2024)
by: Jiang, Kai, et al.
Published: (2024)
GaussianFusion: Gaussian-Based Multi-Sensor Fusion for End-to-End Autonomous Driving
by: Liu, Shuai, et al.
Published: (2025)
by: Liu, Shuai, et al.
Published: (2025)
VLM-E2E: Enhancing End-to-End Autonomous Driving with Multimodal Driver Attention Fusion
by: Liu, Pei, et al.
Published: (2025)
by: Liu, Pei, et al.
Published: (2025)
Physics Inspired Criterion for Pruning-Quantization Joint Learning
by: Xie, Weiying, et al.
Published: (2023)
by: Xie, Weiying, et al.
Published: (2023)
End-to-End 3D Spatiotemporal Perception with Multimodal Fusion and V2X Collaboration
by: Yang, Zhenwei, et al.
Published: (2025)
by: Yang, Zhenwei, et al.
Published: (2025)
An End-to-End Real-World Camera Imaging Pipeline
by: Xu, Kepeng, et al.
Published: (2024)
by: Xu, Kepeng, et al.
Published: (2024)
TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation
by: You, Ling, et al.
Published: (2025)
by: You, Ling, et al.
Published: (2025)
OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models
by: Zhong, Yufeng, et al.
Published: (2026)
by: Zhong, Yufeng, et al.
Published: (2026)
Hyperspectral Mamba for Hyperspectral Object Tracking
by: Gao, Long, et al.
Published: (2025)
by: Gao, Long, et al.
Published: (2025)
ChartE$^{3}$: A Comprehensive Benchmark for End-to-End Chart Editing
by: Li, Shuo, et al.
Published: (2026)
by: Li, Shuo, et al.
Published: (2026)
DA-BEV: Unsupervised Domain Adaptation for Bird's Eye View Perception
by: Jiang, Kai, et al.
Published: (2024)
by: Jiang, Kai, et al.
Published: (2024)
OED: Towards One-stage End-to-End Dynamic Scene Graph Generation
by: Wang, Guan, et al.
Published: (2024)
by: Wang, Guan, et al.
Published: (2024)
Towards Efficient and Effective Multi-Camera Encoding for End-to-End Driving
by: Yang, Jiawei, et al.
Published: (2025)
by: Yang, Jiawei, et al.
Published: (2025)
FusionTrack: End-to-End Multi-Object Tracking in Arbitrary Multi-View Environment
by: Li, Xiaohe, et al.
Published: (2025)
by: Li, Xiaohe, et al.
Published: (2025)
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs
by: Cheng, Dabing, et al.
Published: (2025)
by: Cheng, Dabing, et al.
Published: (2025)
Tracking by Detection and Query: An Efficient End-to-End Framework for Multi-Object Tracking
by: Jia, Shukun, et al.
Published: (2024)
by: Jia, Shukun, et al.
Published: (2024)
Hyperspectral Adapter for Object Tracking based on Hyperspectral Video
by: Gao, Long, et al.
Published: (2025)
by: Gao, Long, et al.
Published: (2025)
E2E-GMNER: End-to-End Generative Grounded Multimodal Named Entity Recognition
by: Zhang, Meng, et al.
Published: (2026)
by: Zhang, Meng, et al.
Published: (2026)
Towards an End-to-End (E2E) Adversarial Learning and Application in the Physical World
by: Biton, Dudi, et al.
Published: (2025)
by: Biton, Dudi, et al.
Published: (2025)
Polar R-CNN: End-to-End Lane Detection with Fewer Anchors
by: Wang, Shengqi, et al.
Published: (2024)
by: Wang, Shengqi, et al.
Published: (2024)
Towards Accurate and Efficient Sub-8-Bit Integer Training
by: Guo, Wenjin, et al.
Published: (2024)
by: Guo, Wenjin, et al.
Published: (2024)
OneVision: An End-to-End Generative Framework for Multi-view E-commerce Vision Search
by: Zheng, Zexin, et al.
Published: (2025)
by: Zheng, Zexin, et al.
Published: (2025)
Beyond Hungarian: Match-Free Supervision for End-to-End Object Detection
by: Qiu, Shoumeng, et al.
Published: (2026)
by: Qiu, Shoumeng, et al.
Published: (2026)
An Effective End-to-End Solution for Multimodal Action Recognition
by: Wang, Songping, et al.
Published: (2025)
by: Wang, Songping, et al.
Published: (2025)
SDformer: Efficient End-to-End Transformer for Depth Completion
by: Qian, Jian, et al.
Published: (2024)
by: Qian, Jian, et al.
Published: (2024)
Towards End-to-End Semi-Supervised Table Detection with Semantic Aligned Matching Transformer
by: Shehzadi, Tahira, et al.
Published: (2024)
by: Shehzadi, Tahira, et al.
Published: (2024)
Similar Items
-
Multi-scale direction-aware SAR object detection network via global information fusion
by: Cao, Mingxiang, et al.
Published: (2023) -
Multimodal Informative ViT: Information Aggregation and Distribution for Hyperspectral and LiDAR Classification
by: Zhang, Jiaqing, et al.
Published: (2024) -
FusionSAM: Visual Multi-Modal Learning with Segment Anything
by: Li, Daixun, et al.
Published: (2024) -
Distribution-aware Interactive Attention Network and Large-scale Cloud Recognition Benchmark on FY-4A Satellite Image
by: Zhang, Jiaqing, et al.
Published: (2024) -
M$^3$amba: CLIP-driven Mamba Model for Multi-modal Remote Sensing Classification
by: Cao, Mingxiang, et al.
Published: (2025)