FOMO-3D: Using Vision Foundation Models for Long-Tailed 3D Object Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Anqi Joyce, Tu, James, Dvornik, Nikita, Li, Enxu, Urtasun, Raquel |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GenAssets: Generating in-the-wild 3D Assets in Latent Space
by: Yang, Ze, et al.
Published: (2026)
by: Yang, Ze, et al.
Published: (2026)
Long-Tailed 3D Detection via Multi-Modal Fusion
by: Ma, Yechi, et al.
Published: (2023)
by: Ma, Yechi, et al.
Published: (2023)
G3R: Gradient Guided Generalizable Reconstruction
by: Chen, Yun, et al.
Published: (2024)
by: Chen, Yun, et al.
Published: (2024)
Inverse++: Vision-Centric 3D Semantic Occupancy Prediction Assisted with 3D Object Detection
by: Ming, Zhenxing, et al.
Published: (2025)
by: Ming, Zhenxing, et al.
Published: (2025)
Bridging Perspectives: Foundation Model Guided BEV Maps for 3D Object Detection and Tracking
by: Käppeler, Markus, et al.
Published: (2025)
by: Käppeler, Markus, et al.
Published: (2025)
Towards Long-Range 3D Object Detection for Autonomous Vehicles
by: Khoche, Ajinkya, et al.
Published: (2023)
by: Khoche, Ajinkya, et al.
Published: (2023)
CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object Detection
by: Kuang, Zhaonian, et al.
Published: (2026)
by: Kuang, Zhaonian, et al.
Published: (2026)
Flux4D: Flow-based Unsupervised 4D Reconstruction
by: Wang, Jingkang, et al.
Published: (2025)
by: Wang, Jingkang, et al.
Published: (2025)
Perspective-Invariant 3D Object Detection
by: Liang, Ao, et al.
Published: (2025)
by: Liang, Ao, et al.
Published: (2025)
Scalable Vision-Based 3D Object Detection and Monocular Depth Estimation for Autonomous Driving
by: Liu, Yuxuan
Published: (2024)
by: Liu, Yuxuan
Published: (2024)
Object-Scene-Camera Decomposition and Recomposition for Data-Efficient Monocular 3D Object Detection
by: Kuang, Zhaonian, et al.
Published: (2026)
by: Kuang, Zhaonian, et al.
Published: (2026)
DeTra: A Unified Model for Object Detection and Trajectory Forecasting
by: Casas, Sergio, et al.
Published: (2024)
by: Casas, Sergio, et al.
Published: (2024)
DIO: Dataset of 3D Mesh Models of Indoor Objects for Robotics and Computer Vision Applications
by: Nimal, Nillan, et al.
Published: (2024)
by: Nimal, Nillan, et al.
Published: (2024)
Rethink 3D Object Detection from Physical World
by: Tanaka, Satoshi, et al.
Published: (2025)
by: Tanaka, Satoshi, et al.
Published: (2025)
Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion
by: Zhang, Lunjun, et al.
Published: (2023)
by: Zhang, Lunjun, et al.
Published: (2023)
DVPE: Divided View Position Embedding for Multi-View 3D Object Detection
by: Wang, Jiasen, et al.
Published: (2024)
by: Wang, Jiasen, et al.
Published: (2024)
Domain Adaptation for Different Sensor Configurations in 3D Object Detection
by: Tanaka, Satoshi, et al.
Published: (2025)
by: Tanaka, Satoshi, et al.
Published: (2025)
Label-Efficient 3D Object Detection For Road-Side Units
by: Dao, Minh-Quan, et al.
Published: (2024)
by: Dao, Minh-Quan, et al.
Published: (2024)
Anyview: Generalizable Indoor 3D Object Detection with Variable Frames
by: Wu, Zhenyu, et al.
Published: (2023)
by: Wu, Zhenyu, et al.
Published: (2023)
Online 3D Scene Reconstruction Using Neural Object Priors
by: Chabal, Thomas, et al.
Published: (2025)
by: Chabal, Thomas, et al.
Published: (2025)
Cross-Level Sensor Fusion with Object Lists via Transformer for 3D Object Detection
by: Liu, Xiangzhong, et al.
Published: (2025)
by: Liu, Xiangzhong, et al.
Published: (2025)
Robust Fusion of Object-Level V2X for Learned 3D Object Detection
by: Ostendorf, Lukas, et al.
Published: (2026)
by: Ostendorf, Lukas, et al.
Published: (2026)
D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
SToRe3D: Sparse Token Relevance in ViTs for Efficient Multi-View 3D Object Detection
by: Papais, Sandro, et al.
Published: (2026)
by: Papais, Sandro, et al.
Published: (2026)
SHOW3D: Capturing Scenes of 3D Hands and Objects in the Wild
by: Rim, Patrick, et al.
Published: (2026)
by: Rim, Patrick, et al.
Published: (2026)
Articulate AnyMesh: Open-Vocabulary 3D Articulated Objects Modeling
by: Qiu, Xiaowen, et al.
Published: (2025)
by: Qiu, Xiaowen, et al.
Published: (2025)
TrackDeform3D: Markerless and Autonomous 3D Keypoint Tracking and Dataset Collection for Deformable Objects
by: Zong, Yeheng, et al.
Published: (2026)
by: Zong, Yeheng, et al.
Published: (2026)
Object and Contact Point Tracking in Demonstrations Using 3D Gaussian Splatting
by: Büttner, Michael, et al.
Published: (2024)
by: Büttner, Michael, et al.
Published: (2024)
Advances in Global Solvers for 3D Vision
by: Zhao, Zhenjun, et al.
Published: (2026)
by: Zhao, Zhenjun, et al.
Published: (2026)
AYDIV: Adaptable Yielding 3D Object Detection via Integrated Contextual Vision Transformer
by: Dam, Tanmoy, et al.
Published: (2024)
by: Dam, Tanmoy, et al.
Published: (2024)
MixSup: Mixed-grained Supervision for Label-efficient LiDAR-based 3D Object Detection
by: Yang, Yuxue, et al.
Published: (2024)
by: Yang, Yuxue, et al.
Published: (2024)
Zero-Shot 3D Visual Grounding from Vision-Language Models
by: Li, Rong, et al.
Published: (2025)
by: Li, Rong, et al.
Published: (2025)
Shelf-Supervised Cross-Modal Pre-Training for 3D Object Detection
by: Khurana, Mehar, et al.
Published: (2024)
by: Khurana, Mehar, et al.
Published: (2024)
LCF3D: A Robust and Real-Time Late-Cascade Fusion Framework for 3D Object Detection in Autonomous Driving
by: Sgaravatti, Carlo, et al.
Published: (2026)
by: Sgaravatti, Carlo, et al.
Published: (2026)
A Comparative Evaluation of Large Vision-Language Models for 2D Object Detection under SOTIF Conditions
by: Zhou, Ji, et al.
Published: (2026)
by: Zhou, Ji, et al.
Published: (2026)
3DGS-CD: 3D Gaussian Splatting-based Change Detection for Physical Object Rearrangement
by: Lu, Ziqi, et al.
Published: (2024)
by: Lu, Ziqi, et al.
Published: (2024)
ZING-3D: Zero-shot Incremental 3D Scene Graphs via Vision-Language Models
by: Saxena, Pranav, et al.
Published: (2025)
by: Saxena, Pranav, et al.
Published: (2025)
LongTail Driving Scenarios with Reasoning Traces: The KITScenes LongTail Dataset
by: Wagner, Royden, et al.
Published: (2026)
by: Wagner, Royden, et al.
Published: (2026)
Cross-Cluster Shifting for Efficient and Effective 3D Object Detection in Autonomous Driving
by: Chen, Zhili, et al.
Published: (2024)
by: Chen, Zhili, et al.
Published: (2024)
Improving Generalization Ability for 3D Object Detection by Learning Sparsity-invariant Features
by: Lu, Hsin-Cheng, et al.
Published: (2025)
by: Lu, Hsin-Cheng, et al.
Published: (2025)
Similar Items
-
GenAssets: Generating in-the-wild 3D Assets in Latent Space
by: Yang, Ze, et al.
Published: (2026) -
Long-Tailed 3D Detection via Multi-Modal Fusion
by: Ma, Yechi, et al.
Published: (2023) -
G3R: Gradient Guided Generalizable Reconstruction
by: Chen, Yun, et al.
Published: (2024) -
Inverse++: Vision-Centric 3D Semantic Occupancy Prediction Assisted with 3D Object Detection
by: Ming, Zhenxing, et al.
Published: (2025) -
Bridging Perspectives: Foundation Model Guided BEV Maps for 3D Object Detection and Tracking
by: Käppeler, Markus, et al.
Published: (2025)