PoIFusion: Multi-Modal 3D Object Detection via Fusion at Points of Interest
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Jiajun, Zhang, Sha, Dayoub, Feras, Ouyang, Wanli, Zhang, Yanyong, Reid, Ian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
by: Deng, Jiajun, et al.
Published: (2025)
by: Deng, Jiajun, et al.
Published: (2025)
HVDistill: Transferring Knowledge from Images to Point Clouds via Unsupervised Hybrid-View Distillation
by: Zhang, Sha, et al.
Published: (2024)
by: Zhang, Sha, et al.
Published: (2024)
Agent3D-Zero: An Agent for Zero-shot 3D Understanding
by: Zhang, Sha, et al.
Published: (2024)
by: Zhang, Sha, et al.
Published: (2024)
RaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera Fusion
by: Chu, Xiaomeng, et al.
Published: (2024)
by: Chu, Xiaomeng, et al.
Published: (2024)
OA-DET3D: Embedding Object Awareness as a General Plug-in for Multi-Camera 3D Object Detection
by: Chu, Xiaomeng, et al.
Published: (2023)
by: Chu, Xiaomeng, et al.
Published: (2023)
Embodied Domain Adaptation for Object Detection
by: Shi, Xiangyu, et al.
Published: (2025)
by: Shi, Xiangyu, et al.
Published: (2025)
RayFormer: Improving Query-Based Multi-Camera 3D Object Detection via Ray-Centric Strategies
by: Chu, Xiaomeng, et al.
Published: (2024)
by: Chu, Xiaomeng, et al.
Published: (2024)
PGOV3D: Open-Vocabulary 3D Semantic Segmentation with Partial-to-Global Curriculum
by: Zhang, Shiqi, et al.
Published: (2025)
by: Zhang, Shiqi, et al.
Published: (2025)
Wasserstein Distance-based Expansion of Low-Density Latent Regions for Unknown Class Detection
by: Mallick, Prakash, et al.
Published: (2024)
by: Mallick, Prakash, et al.
Published: (2024)
SceneEdited: A City-Scale Benchmark for 3D HD Map Updating via Image-Guided Change Detection
by: Lin, Chun-Jung, et al.
Published: (2025)
by: Lin, Chun-Jung, et al.
Published: (2025)
ObjectReact: Learning Object-Relative Control for Visual Navigation
by: Garg, Sourav, et al.
Published: (2025)
by: Garg, Sourav, et al.
Published: (2025)
Improving Online Source-free Domain Adaptation for Object Detection by Unsupervised Data Acquisition
by: Shi, Xiangyu, et al.
Published: (2023)
by: Shi, Xiangyu, et al.
Published: (2023)
Detecting Precise Hand Touch Moments in Egocentric Video
by: Nguyen, Huy Anh, et al.
Published: (2026)
by: Nguyen, Huy Anh, et al.
Published: (2026)
To Ask or Not to Ask? Detecting Absence of Information in Vision and Language Navigation
by: Abraham, Savitha Sam, et al.
Published: (2024)
by: Abraham, Savitha Sam, et al.
Published: (2024)
Progressive Modality Cooperation for Multi-Modality Domain Adaptation
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
Temporal Attention for Cross-View Sequential Image Localization
by: Yuan, Dong, et al.
Published: (2024)
by: Yuan, Dong, et al.
Published: (2024)
TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals
by: Podgorski, Stefan, et al.
Published: (2025)
by: Podgorski, Stefan, et al.
Published: (2025)
Segment Beyond View: Handling Partially Missing Modality for Audio-Visual Semantic Segmentation
by: Wu, Renjie, et al.
Published: (2023)
by: Wu, Renjie, et al.
Published: (2023)
BEVDilation: LiDAR-Centric Multi-Modal Fusion for 3D Object Detection
by: Zhang, Guowen, et al.
Published: (2025)
by: Zhang, Guowen, et al.
Published: (2025)
RoboFusion: Towards Robust Multi-Modal 3D Object Detection via SAM
by: Song, Ziying, et al.
Published: (2024)
by: Song, Ziying, et al.
Published: (2024)
A Physical Agentic Loop for Language-Guided Grasping with Execution-State Monitoring
by: Wang, Wenze, et al.
Published: (2026)
by: Wang, Wenze, et al.
Published: (2026)
Progressive Multi-Modal Fusion for Robust 3D Object Detection
by: Mohan, Rohit, et al.
Published: (2024)
by: Mohan, Rohit, et al.
Published: (2024)
Multi-scale Feature Fusion with Point Pyramid for 3D Object Detection
by: Lu, Weihao, et al.
Published: (2024)
by: Lu, Weihao, et al.
Published: (2024)
3D Object Detection from Images for Autonomous Driving: A Survey
by: Ma, Xinzhu, et al.
Published: (2022)
by: Ma, Xinzhu, et al.
Published: (2022)
KITE: Keyframe-Indexed Tokenized Evidence for VLM-Based Robot Failure Analysis
by: Hosseinzadeh, Mehdi, et al.
Published: (2026)
by: Hosseinzadeh, Mehdi, et al.
Published: (2026)
Learning Geometry-Guided Depth via Projective Modeling for Monocular 3D Object Detection
by: Zhang, Yinmin, et al.
Published: (2021)
by: Zhang, Yinmin, et al.
Published: (2021)
Robust Scene Change Detection Using Visual Foundation Models and Cross-Attention Mechanisms
by: Lin, Chun-Jung, et al.
Published: (2024)
by: Lin, Chun-Jung, et al.
Published: (2024)
RoboHop: Segment-based Topological Map Representation for Open-World Visual Navigation
by: Garg, Sourav, et al.
Published: (2024)
by: Garg, Sourav, et al.
Published: (2024)
QueryAdapter: Rapid Adaptation of Vision-Language Models in Response to Natural Language Queries
by: Chapman, Nicolas Harvey, et al.
Published: (2025)
by: Chapman, Nicolas Harvey, et al.
Published: (2025)
Fusion is Not Enough: Single Modal Attacks on Fusion Models for 3D Object Detection
by: Cheng, Zhiyuan, et al.
Published: (2023)
by: Cheng, Zhiyuan, et al.
Published: (2023)
VoxelNextFusion: A Simple, Unified and Effective Voxel Fusion Framework for Multi-Modal 3D Object Detection
by: Song, Ziying, et al.
Published: (2024)
by: Song, Ziying, et al.
Published: (2024)
DGFusion: Dual-guided Fusion for Robust Multi-Modal 3D Object Detection
by: Jia, Feiyang, et al.
Published: (2025)
by: Jia, Feiyang, et al.
Published: (2025)
Enhancing Embodied Object Detection through Language-Image Pre-training and Implicit Object Memory
by: Chapman, Nicolas Harvey, et al.
Published: (2024)
by: Chapman, Nicolas Harvey, et al.
Published: (2024)
AIMC-Spec: A Benchmark Dataset for Automatic Intrapulse Modulation Classification under Variable Noise Conditions
by: Cocks, Sebastian L., et al.
Published: (2026)
by: Cocks, Sebastian L., et al.
Published: (2026)
MultiCorrupt: A Multi-Modal Robustness Dataset and Benchmark of LiDAR-Camera Fusion for 3D Object Detection
by: Beemelmanns, Till, et al.
Published: (2024)
by: Beemelmanns, Till, et al.
Published: (2024)
SpatialSplat: Efficient Semantic 3D from Sparse Unposed Images
by: Sheng, Yu, et al.
Published: (2025)
by: Sheng, Yu, et al.
Published: (2025)
CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection
by: Wu, Yuchen, et al.
Published: (2026)
by: Wu, Yuchen, et al.
Published: (2026)
Long-Tailed 3D Detection via Multi-Modal Fusion
by: Ma, Yechi, et al.
Published: (2023)
by: Ma, Yechi, et al.
Published: (2023)
Semi-supervised 3D Object Detection with PatchTeacher and PillarMix
by: Wu, Xiaopei, et al.
Published: (2024)
by: Wu, Xiaopei, et al.
Published: (2024)
PIGEON: VLM-Driven Object Navigation via Points of Interest Selection
by: Peng, Cheng, et al.
Published: (2025)
by: Peng, Cheng, et al.
Published: (2025)
Similar Items
-
3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
by: Deng, Jiajun, et al.
Published: (2025) -
HVDistill: Transferring Knowledge from Images to Point Clouds via Unsupervised Hybrid-View Distillation
by: Zhang, Sha, et al.
Published: (2024) -
Agent3D-Zero: An Agent for Zero-shot 3D Understanding
by: Zhang, Sha, et al.
Published: (2024) -
RaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera Fusion
by: Chu, Xiaomeng, et al.
Published: (2024) -
OA-DET3D: Embedding Object Awareness as a General Plug-in for Multi-Camera 3D Object Detection
by: Chu, Xiaomeng, et al.
Published: (2023)