CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Tianjiao, Li, Xinzhuo, Shen, Yifan, Liu, Yuanzhe, Lourentzou, Ismini |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising
by: Yu, Tianjiao, et al.
Published: (2026)
by: Yu, Tianjiao, et al.
Published: (2026)
Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination
by: Li, Xinzhuo, et al.
Published: (2025)
by: Li, Xinzhuo, et al.
Published: (2025)
3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding
by: Ogunleye, Makanjuola, et al.
Published: (2026)
by: Ogunleye, Makanjuola, et al.
Published: (2026)
Part$^{2}$GS: Part-aware Modeling of Articulated Objects using 3D Gaussian Splatting
by: Yu, Tianjiao, et al.
Published: (2025)
by: Yu, Tianjiao, et al.
Published: (2025)
CALICO: Part-Focused Semantic Co-Segmentation with Large Vision-Language Models
by: Nguyen, Kiet A., et al.
Published: (2024)
by: Nguyen, Kiet A., et al.
Published: (2024)
PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation
by: Wahed, Muntasir, et al.
Published: (2024)
by: Wahed, Muntasir, et al.
Published: (2024)
Commonsense for Zero-Shot Natural Language Video Localization
by: Holla, Meghana, et al.
Published: (2023)
by: Holla, Meghana, et al.
Published: (2023)
ELBA: Learning by Asking for Embodied Visual Navigation and Task Completion
by: Shen, Ying, et al.
Published: (2023)
by: Shen, Ying, et al.
Published: (2023)
VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
by: Yuan, Haoran, et al.
Published: (2026)
by: Yuan, Haoran, et al.
Published: (2026)
Toward Cognitive Supersensing in Multimodal Large Language Model
by: Li, Boyi, et al.
Published: (2026)
by: Li, Boyi, et al.
Published: (2026)
BIP3D: Bridging 2D Images and 3D Perception for Embodied Intelligence
by: Lin, Xuewu, et al.
Published: (2024)
by: Lin, Xuewu, et al.
Published: (2024)
3D-LFM: Lifting Foundation Model
by: Dabhi, Mosam, et al.
Published: (2023)
by: Dabhi, Mosam, et al.
Published: (2023)
Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics
by: Shen, Ying, et al.
Published: (2026)
by: Shen, Ying, et al.
Published: (2026)
Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single Video
by: Yao, David Yifan, et al.
Published: (2025)
by: Yao, David Yifan, et al.
Published: (2025)
Pixels to Play: A Foundation Model for 3D Gameplay
by: Yue, Yuguang, et al.
Published: (2025)
by: Yue, Yuguang, et al.
Published: (2025)
CoReEcho: Continuous Representation Learning for 2D+time Echocardiography Analysis
by: Maani, Fadillah Adamsyah, et al.
Published: (2024)
by: Maani, Fadillah Adamsyah, et al.
Published: (2024)
Decipher-MR: A Vision-Language Foundation Model for 3D MRI Representations
by: Yang, Zhijian, et al.
Published: (2025)
by: Yang, Zhijian, et al.
Published: (2025)
Towards Fair Medical AI: Adversarial Debiasing of 3D CT Foundation Embeddings
by: Zheng, Guangyao, et al.
Published: (2025)
by: Zheng, Guangyao, et al.
Published: (2025)
Flame3D: Zero-shot Compositional Reasoning of 3D Scenes with Agentic Language Models
by: Bharadwaj, Sagar, et al.
Published: (2026)
by: Bharadwaj, Sagar, et al.
Published: (2026)
Instruct 4D-to-4D: Editing 4D Scenes as Pseudo-3D Scenes Using 2D Diffusion
by: Mou, Linzhan, et al.
Published: (2024)
by: Mou, Linzhan, et al.
Published: (2024)
Knowledge Transfer Scaling Laws for 3D Medical Imaging
by: Lee, Ho Hin, et al.
Published: (2026)
by: Lee, Ho Hin, et al.
Published: (2026)
Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
by: Man, Yunze, et al.
Published: (2024)
by: Man, Yunze, et al.
Published: (2024)
3D ReX: Causal Explanations in 3D Neuroimaging Classification
by: Navaratnarajah, Melane, et al.
Published: (2025)
by: Navaratnarajah, Melane, et al.
Published: (2025)
Situational Awareness Matters in 3D Vision Language Reasoning
by: Man, Yunze, et al.
Published: (2024)
by: Man, Yunze, et al.
Published: (2024)
Research on the Spatial Data Intelligent Foundation Model
by: Wang, Shaohua, et al.
Published: (2024)
by: Wang, Shaohua, et al.
Published: (2024)
2D-3D Interlaced Transformer for Point Cloud Segmentation with Scene-Level Supervision
by: Yang, Cheng-Kun, et al.
Published: (2023)
by: Yang, Cheng-Kun, et al.
Published: (2023)
Demographic Predictability in 3D CT Foundation Embeddings
by: Zheng, Guangyao, et al.
Published: (2024)
by: Zheng, Guangyao, et al.
Published: (2024)
EgoForge: Goal-Directed Egocentric World Simulator
by: Shen, Yifan, et al.
Published: (2026)
by: Shen, Yifan, et al.
Published: (2026)
Revisiting CLIP: Efficient Alignment of 3D MRI and Tabular Data using Domain-Specific Foundation Models
by: Petersen, Jakob Krogh, et al.
Published: (2025)
by: Petersen, Jakob Krogh, et al.
Published: (2025)
TAPVid-3D: A Benchmark for Tracking Any Point in 3D
by: Koppula, Skanda, et al.
Published: (2024)
by: Koppula, Skanda, et al.
Published: (2024)
DF-Mamba: Deformable State Space Modeling for 3D Hand Pose Estimation in Interactions
by: Zhou, Yifan, et al.
Published: (2025)
by: Zhou, Yifan, et al.
Published: (2025)
Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models
by: Zhou, Shengli, et al.
Published: (2026)
by: Zhou, Shengli, et al.
Published: (2026)
GSNeRF: Generalizable Semantic Neural Radiance Fields with Enhanced 3D Scene Understanding
by: Chou, Zi-Ting, et al.
Published: (2024)
by: Chou, Zi-Ting, et al.
Published: (2024)
3D Reconstruction of Objects in Hands without Real World 3D Supervision
by: Prakash, Aditya, et al.
Published: (2023)
by: Prakash, Aditya, et al.
Published: (2023)
Grounding Foundational Vision Models with 3D Human Poses for Robust Action Recognition
by: Babey, Nicholas, et al.
Published: (2025)
by: Babey, Nicholas, et al.
Published: (2025)
ConsistDreamer: 3D-Consistent 2D Diffusion for High-Fidelity Scene Editing
by: Chen, Jun-Kun, et al.
Published: (2024)
by: Chen, Jun-Kun, et al.
Published: (2024)
IM-3D: Iterative Multiview Diffusion and Reconstruction for High-Quality 3D Generation
by: Melas-Kyriazi, Luke, et al.
Published: (2024)
by: Melas-Kyriazi, Luke, et al.
Published: (2024)
RQR3D: Reparametrizing the regression targets for BEV-based 3D object detection
by: Kilinc, Ozsel, et al.
Published: (2025)
by: Kilinc, Ozsel, et al.
Published: (2025)
Collaborating Foundation Models for Domain Generalized Semantic Segmentation
by: Benigmim, Yasser, et al.
Published: (2023)
by: Benigmim, Yasser, et al.
Published: (2023)
OpenSUN3D: 1st Workshop Challenge on Open-Vocabulary 3D Scene Understanding
by: Engelmann, Francis, et al.
Published: (2024)
by: Engelmann, Francis, et al.
Published: (2024)
Similar Items
-
DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising
by: Yu, Tianjiao, et al.
Published: (2026) -
Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination
by: Li, Xinzhuo, et al.
Published: (2025) -
3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding
by: Ogunleye, Makanjuola, et al.
Published: (2026) -
Part$^{2}$GS: Part-aware Modeling of Articulated Objects using 3D Gaussian Splatting
by: Yu, Tianjiao, et al.
Published: (2025) -
CALICO: Part-Focused Semantic Co-Segmentation with Large Vision-Language Models
by: Nguyen, Kiet A., et al.
Published: (2024)