BIP3D: Bridging 2D Images and 3D Perception for Embodied Intelligence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Xuewu, Lin, Tianwei, Huang, Lichao, Xie, Hongyu, Su, Zhizhong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SEM: Enhancing Spatial Understanding for Robust Robot Manipulation
von: Lin, Xuewu, et al.
Veröffentlicht: (2025)
von: Lin, Xuewu, et al.
Veröffentlicht: (2025)
EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
von: Wu, Yuqi, et al.
Veröffentlicht: (2024)
von: Wu, Yuqi, et al.
Veröffentlicht: (2024)
An Embodied Generalist Agent in 3D World
von: Huang, Jiangyong, et al.
Veröffentlicht: (2023)
von: Huang, Jiangyong, et al.
Veröffentlicht: (2023)
PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI
von: Yang, Yandan, et al.
Veröffentlicht: (2024)
von: Yang, Yandan, et al.
Veröffentlicht: (2024)
CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
von: Yu, Tianjiao, et al.
Veröffentlicht: (2025)
von: Yu, Tianjiao, et al.
Veröffentlicht: (2025)
3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding
von: Ogunleye, Makanjuola, et al.
Veröffentlicht: (2026)
von: Ogunleye, Makanjuola, et al.
Veröffentlicht: (2026)
2D-3D Interlaced Transformer for Point Cloud Segmentation with Scene-Level Supervision
von: Yang, Cheng-Kun, et al.
Veröffentlicht: (2023)
von: Yang, Cheng-Kun, et al.
Veröffentlicht: (2023)
landmarker: a Toolkit for Anatomical Landmark Localization in 2D/3D Images
von: Jonkers, Jef, et al.
Veröffentlicht: (2025)
von: Jonkers, Jef, et al.
Veröffentlicht: (2025)
RoboLayout: Differentiable 3D Scene Generation for Embodied Agents
von: Shamsaddinlou, Ali
Veröffentlicht: (2026)
von: Shamsaddinlou, Ali
Veröffentlicht: (2026)
SPA: 3D Spatial-Awareness Enables Effective Embodied Representation
von: Zhu, Haoyi, et al.
Veröffentlicht: (2024)
von: Zhu, Haoyi, et al.
Veröffentlicht: (2024)
Bridging the Gap between 2D and 3D Visual Question Answering: A Fusion Approach for 3D VQA
von: Mo, Wentao, et al.
Veröffentlicht: (2024)
von: Mo, Wentao, et al.
Veröffentlicht: (2024)
Semi-Supervised 3D Medical Segmentation from 2D Natural Images Pretrained Model
von: Yeung, Pak-Hei, et al.
Veröffentlicht: (2025)
von: Yeung, Pak-Hei, et al.
Veröffentlicht: (2025)
ICE-G: Image Conditional Editing of 3D Gaussian Splats
von: Jaganathan, Vishnu, et al.
Veröffentlicht: (2024)
von: Jaganathan, Vishnu, et al.
Veröffentlicht: (2024)
Uncertainty-aware Diffusion and Reinforcement Learning for Joint Plane Localization and Anomaly Diagnosis in 3D Ultrasound
von: Huang, Yuhao, et al.
Veröffentlicht: (2025)
von: Huang, Yuhao, et al.
Veröffentlicht: (2025)
GrabS: Generative Embodied Agent for 3D Object Segmentation without Scene Supervision
von: Zhang, Zihui, et al.
Veröffentlicht: (2025)
von: Zhang, Zihui, et al.
Veröffentlicht: (2025)
Instant3D: Instant Text-to-3D Generation
von: Li, Ming, et al.
Veröffentlicht: (2023)
von: Li, Ming, et al.
Veröffentlicht: (2023)
STELLAR: Scaling 3D Perception Large Models for Autonomous Driving
von: Li, Yingwei, et al.
Veröffentlicht: (2026)
von: Li, Yingwei, et al.
Veröffentlicht: (2026)
Instruct 4D-to-4D: Editing 4D Scenes as Pseudo-3D Scenes Using 2D Diffusion
von: Mou, Linzhan, et al.
Veröffentlicht: (2024)
von: Mou, Linzhan, et al.
Veröffentlicht: (2024)
From 2D to 3D, Deep Learning-based Shape Reconstruction in Magnetic Resonance Imaging: A Review
von: McMillian, Emma, et al.
Veröffentlicht: (2025)
von: McMillian, Emma, et al.
Veröffentlicht: (2025)
AmaraSpatial-10K: A Spatially and Semantically Aligned 3D Dataset for Spatial Computing and Embodied AI
von: Salehi, Mohammad Sadegh, et al.
Veröffentlicht: (2026)
von: Salehi, Mohammad Sadegh, et al.
Veröffentlicht: (2026)
Knowledge Transfer Scaling Laws for 3D Medical Imaging
von: Lee, Ho Hin, et al.
Veröffentlicht: (2026)
von: Lee, Ho Hin, et al.
Veröffentlicht: (2026)
TAPVid-3D: A Benchmark for Tracking Any Point in 3D
von: Koppula, Skanda, et al.
Veröffentlicht: (2024)
von: Koppula, Skanda, et al.
Veröffentlicht: (2024)
SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters
von: Jiang, Jianping, et al.
Veröffentlicht: (2024)
von: Jiang, Jianping, et al.
Veröffentlicht: (2024)
Learning with Noisy Ground Truth: From 2D Classification to 3D Reconstruction
von: Lu, Yangdi, et al.
Veröffentlicht: (2024)
von: Lu, Yangdi, et al.
Veröffentlicht: (2024)
3D Reconstruction of Objects in Hands without Real World 3D Supervision
von: Prakash, Aditya, et al.
Veröffentlicht: (2023)
von: Prakash, Aditya, et al.
Veröffentlicht: (2023)
Bimanual 3D Hand Motion and Articulation Forecasting in Everyday Images
von: Prakash, Aditya, et al.
Veröffentlicht: (2025)
von: Prakash, Aditya, et al.
Veröffentlicht: (2025)
RadarOcc: Robust 3D Occupancy Prediction with 4D Imaging Radar
von: Ding, Fangqiang, et al.
Veröffentlicht: (2024)
von: Ding, Fangqiang, et al.
Veröffentlicht: (2024)
Salsa as a Nonverbal Embodied Language -- The CoMPAS3D Dataset and Benchmarks
von: Burkanova, Bermet, et al.
Veröffentlicht: (2025)
von: Burkanova, Bermet, et al.
Veröffentlicht: (2025)
Solving 3D Inverse Problems using Pre-trained 2D Diffusion Models
von: Chung, Hyungjin, et al.
Veröffentlicht: (2022)
von: Chung, Hyungjin, et al.
Veröffentlicht: (2022)
ConsistDreamer: 3D-Consistent 2D Diffusion for High-Fidelity Scene Editing
von: Chen, Jun-Kun, et al.
Veröffentlicht: (2024)
von: Chen, Jun-Kun, et al.
Veröffentlicht: (2024)
RadSAM: Segmenting 3D radiological images with a 2D promptable model
von: Khlaut, Julien, et al.
Veröffentlicht: (2025)
von: Khlaut, Julien, et al.
Veröffentlicht: (2025)
IM-3D: Iterative Multiview Diffusion and Reconstruction for High-Quality 3D Generation
von: Melas-Kyriazi, Luke, et al.
Veröffentlicht: (2024)
von: Melas-Kyriazi, Luke, et al.
Veröffentlicht: (2024)
RQR3D: Reparametrizing the regression targets for BEV-based 3D object detection
von: Kilinc, Ozsel, et al.
Veröffentlicht: (2025)
von: Kilinc, Ozsel, et al.
Veröffentlicht: (2025)
OpenSUN3D: 1st Workshop Challenge on Open-Vocabulary 3D Scene Understanding
von: Engelmann, Francis, et al.
Veröffentlicht: (2024)
von: Engelmann, Francis, et al.
Veröffentlicht: (2024)
3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
von: Hu, Wenbo, et al.
Veröffentlicht: (2025)
von: Hu, Wenbo, et al.
Veröffentlicht: (2025)
ODIN: A Single Model for 2D and 3D Segmentation
von: Jain, Ayush, et al.
Veröffentlicht: (2024)
von: Jain, Ayush, et al.
Veröffentlicht: (2024)
MuM: Multi-View Masked Image Modeling for 3D Vision
von: Nordström, David, et al.
Veröffentlicht: (2025)
von: Nordström, David, et al.
Veröffentlicht: (2025)
A Recipe for Generating 3D Worlds From a Single Image
von: Schwarz, Katja, et al.
Veröffentlicht: (2025)
von: Schwarz, Katja, et al.
Veröffentlicht: (2025)
GaussianFormer-2: Probabilistic Gaussian Superposition for Efficient 3D Occupancy Prediction
von: Huang, Yuanhui, et al.
Veröffentlicht: (2024)
von: Huang, Yuanhui, et al.
Veröffentlicht: (2024)
SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards
von: Batra, Hunar, et al.
Veröffentlicht: (2025)
von: Batra, Hunar, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SEM: Enhancing Spatial Understanding for Robust Robot Manipulation
von: Lin, Xuewu, et al.
Veröffentlicht: (2025) -
EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
von: Wu, Yuqi, et al.
Veröffentlicht: (2024) -
An Embodied Generalist Agent in 3D World
von: Huang, Jiangyong, et al.
Veröffentlicht: (2023) -
PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI
von: Yang, Yandan, et al.
Veröffentlicht: (2024) -
CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
von: Yu, Tianjiao, et al.
Veröffentlicht: (2025)