Robust Multimodal 3D Object Detection via Modality-Agnostic Decoding and Proximity-based Modality Ensemble
Fuente:
arXiv
Saved in:
| Main Authors: | Cha, Juhan, Joo, Minseok, Park, Jihwan, Lee, Sanghyeok, Kim, Injae, Kim, Hyunwoo J. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
F4Splat: Feed-Forward Predictive Densification for Feed-Forward 3D Gaussian Splatting
by: Kim, Injae, et al.
Published: (2026)
by: Kim, Injae, et al.
Published: (2026)
Efficient multi-view training for 3D Gaussian Splatting
by: Choi, Minhyuk, et al.
Published: (2025)
by: Choi, Minhyuk, et al.
Published: (2025)
Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
by: Park, Jihwan, et al.
Published: (2025)
by: Park, Jihwan, et al.
Published: (2025)
Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality Generation
by: Park, Dogyun, et al.
Published: (2025)
by: Park, Dogyun, et al.
Published: (2025)
EfficientViM: Efficient Vision Mamba with Hidden State Mixer based State Space Duality
by: Lee, Sanghyeok, et al.
Published: (2024)
by: Lee, Sanghyeok, et al.
Published: (2024)
RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction Detection
by: Park, Jihwan, et al.
Published: (2026)
by: Park, Jihwan, et al.
Published: (2026)
Groupwise Query Specialization and Quality-Aware Multi-Assignment for Transformer-based Visual Relationship Detection
by: Kim, Jongha, et al.
Published: (2024)
by: Kim, Jongha, et al.
Published: (2024)
Retrieve What's Missing: Coverage-Maximizing Retrieval for Consistent Long Video Generation
by: Joo, Minseok, et al.
Published: (2026)
by: Joo, Minseok, et al.
Published: (2026)
HiddenObject: Modality-Agnostic Fusion for Multimodal Hidden Object Detection
by: Song, Harris, et al.
Published: (2025)
by: Song, Harris, et al.
Published: (2025)
Multi-criteria Token Fusion with One-step-ahead Attention for Efficient Vision Transformers
by: Lee, Sanghyeok, et al.
Published: (2024)
by: Lee, Sanghyeok, et al.
Published: (2024)
CMTM: Cross-Modal Token Modulation for Unsupervised Video Object Segmentation
by: Jeon, Inseok, et al.
Published: (2026)
by: Jeon, Inseok, et al.
Published: (2026)
Modality-Agnostic Prompt Learning for Multi-Modal Camouflaged Object Detection
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
EVT: Efficient View Transformation for Multi-Modal 3D Object Detection
by: Lee, Yongjin, et al.
Published: (2024)
by: Lee, Yongjin, et al.
Published: (2024)
Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI Detection
by: Yang, Chanhyeong, et al.
Published: (2025)
by: Yang, Chanhyeong, et al.
Published: (2025)
CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object Detection
by: Chang, Gyusam, et al.
Published: (2024)
by: Chang, Gyusam, et al.
Published: (2024)
MoE-GRPO: Optimizing Mixture-of-Experts via Reinforcement Learning in Vision-Language Models
by: Ko, Dohwan, et al.
Published: (2026)
by: Ko, Dohwan, et al.
Published: (2026)
Dynamic Full-body Motion Agent with Object Interaction via Blending Pre-trained Modular Controllers
by: Nam, Sanghyeok, et al.
Published: (2026)
by: Nam, Sanghyeok, et al.
Published: (2026)
ModalPatch: A Plug-and-Play Module for Robust Multi-Modal 3D Object Detection under Modality Drop
by: Li, Shuangzhi, et al.
Published: (2026)
by: Li, Shuangzhi, et al.
Published: (2026)
TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing
by: Kim, Jongha, et al.
Published: (2025)
by: Kim, Jongha, et al.
Published: (2025)
Progressive Multi-Modal Fusion for Robust 3D Object Detection
by: Mohan, Rohit, et al.
Published: (2024)
by: Mohan, Rohit, et al.
Published: (2024)
SpatialMosaic: A Multiview VLM Dataset for Partial Visibility
by: Lee, Kanghee, et al.
Published: (2025)
by: Lee, Kanghee, et al.
Published: (2025)
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
by: Park, Kyu Ri, et al.
Published: (2024)
by: Park, Kyu Ri, et al.
Published: (2024)
Panopticus: Omnidirectional 3D Object Detection on Resource-constrained Edge Devices
by: Lee, Jeho, et al.
Published: (2024)
by: Lee, Jeho, et al.
Published: (2024)
DocPrune:Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning
by: Choi, Joonmyung, et al.
Published: (2026)
by: Choi, Joonmyung, et al.
Published: (2026)
RoboFusion: Towards Robust Multi-Modal 3D Object Detection via SAM
by: Song, Ziying, et al.
Published: (2024)
by: Song, Ziying, et al.
Published: (2024)
Super-class guided Transformer for Zero-Shot Attribute Classification
by: Kim, Sehyung, et al.
Published: (2025)
by: Kim, Sehyung, et al.
Published: (2025)
Modality-Agnostic fMRI Decoding of Vision and Language
by: Nikolaus, Mitja, et al.
Published: (2024)
by: Nikolaus, Mitja, et al.
Published: (2024)
Multi-Modal Decouple and Recouple Network for Robust 3D Object Detection
by: Ding, Rui, et al.
Published: (2026)
by: Ding, Rui, et al.
Published: (2026)
PEGASUS: Personalized Generative 3D Avatars with Composable Attributes
by: Cha, Hyunsoo, et al.
Published: (2024)
by: Cha, Hyunsoo, et al.
Published: (2024)
Multi-Modal Guided Multi-Source Domain Adaptation for Object Detection
by: Lee, Sangin, et al.
Published: (2026)
by: Lee, Sangin, et al.
Published: (2026)
Contrast-Guided Cross-Modal Distillation for Thermal Object Detection
by: Kim, SiWoo, et al.
Published: (2025)
by: Kim, SiWoo, et al.
Published: (2025)
MiPa: Mixed Patch Infrared-Visible Modality Agnostic Object Detection
by: Medeiros, Heitor R., et al.
Published: (2024)
by: Medeiros, Heitor R., et al.
Published: (2024)
Generative Subgraph Retrieval for Knowledge Graph-Grounded Dialog Generation
by: Park, Jinyoung, et al.
Published: (2024)
by: Park, Jinyoung, et al.
Published: (2024)
Layer-Wise Modality Decomposition for Interpretable Multimodal Sensor Fusion
by: Park, Jaehyun, et al.
Published: (2025)
by: Park, Jaehyun, et al.
Published: (2025)
Retrieval-Augmented Open-Vocabulary Object Detection
by: Kim, Jooyeon, et al.
Published: (2024)
by: Kim, Jooyeon, et al.
Published: (2024)
Adaptive LiDAR Scanning: Harnessing Temporal Cues for Efficient 3D Object Detection via Multi-Modal Fusion
by: Shoouri, Sara, et al.
Published: (2025)
by: Shoouri, Sara, et al.
Published: (2025)
Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations
by: Kim, Jeonghyeon, et al.
Published: (2025)
by: Kim, Jeonghyeon, et al.
Published: (2025)
Modality-Agnostic Style Transfer for Holistic Feature Imputation
by: Baek, Seunghun, et al.
Published: (2025)
by: Baek, Seunghun, et al.
Published: (2025)
DGFusion: Dual-guided Fusion for Robust Multi-Modal 3D Object Detection
by: Jia, Feiyang, et al.
Published: (2025)
by: Jia, Feiyang, et al.
Published: (2025)
Efficient Test-Time Optimization for Depth Completion via Low-Rank Decoder Adaptation
by: Seo, Minseok, et al.
Published: (2026)
by: Seo, Minseok, et al.
Published: (2026)
Similar Items
-
F4Splat: Feed-Forward Predictive Densification for Feed-Forward 3D Gaussian Splatting
by: Kim, Injae, et al.
Published: (2026) -
Efficient multi-view training for 3D Gaussian Splatting
by: Choi, Minhyuk, et al.
Published: (2025) -
Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
by: Park, Jihwan, et al.
Published: (2025) -
Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality Generation
by: Park, Dogyun, et al.
Published: (2025) -
EfficientViM: Efficient Vision Mamba with Hidden State Mixer based State Space Duality
by: Lee, Sanghyeok, et al.
Published: (2024)