Reasoning3D -- Grounding and Reasoning in 3D: Fine-Grained Zero-Shot Open-Vocabulary 3D Reasoning Part Segmentation via Large Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Tianrun, Yu, Chunan, Li, Jing, Zhang, Jianqi, Zhu, Lanyun, Ji, Deyi, Zhang, Yong, Zang, Ying, Li, Zejian, Sun, Lingyun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
xLSTM-UNet can be an Effective 2D & 3D Medical Image Segmentation Backbone with Vision-LSTM (ViL) better than its Mamba Counterpart
by: Chen, Tianrun, et al.
Published: (2024)
by: Chen, Tianrun, et al.
Published: (2024)
Adversarial Unsupervised Domain Adaptation for 3D Semantic Segmentation with 2D Image Fusion of Dense Depth
by: Xindan Zhang, et al.
Published: (2024)
by: Xindan Zhang, et al.
Published: (2024)
SAM2-Adapter: Evaluating & Adapting Segment Anything 2 in Downstream Tasks: Camouflage, Shadow, Medical Image Segmentation, and More
by: Chen, Tianrun, et al.
Published: (2024)
by: Chen, Tianrun, et al.
Published: (2024)
Taming Video Models for 3D and 4D Generation via Zero-Shot Camera Control
by: Song, Chenxi, et al.
Published: (2025)
by: Song, Chenxi, et al.
Published: (2025)
AdaToken-3D: Dynamic Spatial Gating for Efficient 3D Large Multimodal-Models Reasoning
by: Zhang, Kai, et al.
Published: (2025)
by: Zhang, Kai, et al.
Published: (2025)
Chirpy3D: Part-Aware Multi-View Diffusion for Creative Fine-Grained Object Generation
by: Ng, Kam Woh, et al.
Published: (2025)
by: Ng, Kam Woh, et al.
Published: (2025)
BoxFusion: Reconstruction‐Free Open‐Vocabulary 3D Object Detection via Real‐Time Multi‐View Box Fusion
by: Yuqing Lan, et al.
Published: (2025)
by: Yuqing Lan, et al.
Published: (2025)
Split&Splat: Zero-Shot Panoptic Segmentation via Explicit Instance Modeling and 3D Gaussian Splatting
by: Monchieri, Leonardo, et al.
Published: (2026)
by: Monchieri, Leonardo, et al.
Published: (2026)
Syllables to Scenes: Literary-Guided Free-Viewpoint 3D Scene Synthesis from Japanese Haiku
by: Yu, Chunan, et al.
Published: (2025)
by: Yu, Chunan, et al.
Published: (2025)
From Air to Wear: Personalized 3D Digital Fashion with AR/VR Immersive 3D Sketching
by: Zang, Ying, et al.
Published: (2025)
by: Zang, Ying, et al.
Published: (2025)
ZeroScene: A Zero-Shot Framework for 3D Scene Generation from a Single Image and Controllable Texture Editing
by: Tang, Xiang, et al.
Published: (2025)
by: Tang, Xiang, et al.
Published: (2025)
Multimodal 3D Few‐Shot Classification via Gaussian Mixture Discriminant Analysis
by: Yiqi Wu, et al.
Published: (2025)
by: Yiqi Wu, et al.
Published: (2025)
GS‐Octree: Octree‐based 3D Gaussian Splatting for Robust Object‐level 3D Reconstruction Under Strong Lighting
by: J. Li, et al.
Published: (2024)
by: J. Li, et al.
Published: (2024)
Contrastive Multi-Modal Hypergraph Reasoning for 3D Crowd Mesh Recovery
by: Sun, Minghao, et al.
Published: (2026)
by: Sun, Minghao, et al.
Published: (2026)
PartUV: Part-Based UV Unwrapping of 3D Meshes
by: Wang, Zhaoning, et al.
Published: (2025)
by: Wang, Zhaoning, et al.
Published: (2025)
FSH3D: 3D Representation via Fibonacci Spherical Harmonics
by: Zikuan Li, et al.
Published: (2024)
by: Zikuan Li, et al.
Published: (2024)
ZeroHSI: Zero-Shot 4D Human-Scene Interaction by Video Generation
by: Li, Hongjie, et al.
Published: (2024)
by: Li, Hongjie, et al.
Published: (2024)
AnyHome: Open-Vocabulary Generation of Structured and Textured 3D Homes
by: Fu, Rao, et al.
Published: (2023)
by: Fu, Rao, et al.
Published: (2023)
Joint Deblurring and 3D Reconstruction for Macrophotography
by: Yifan Zhao, et al.
Published: (2025)
by: Yifan Zhao, et al.
Published: (2025)
Back to 3D: Few-Shot 3D Keypoint Detection with Back-Projected 2D Features
by: Wimmer, Thomas, et al.
Published: (2023)
by: Wimmer, Thomas, et al.
Published: (2023)
Capture, Canonicalize, Splat: Zero-Shot 3D Gaussian Avatars from Unstructured Phone Images
by: Garbin, Emanuel, et al.
Published: (2025)
by: Garbin, Emanuel, et al.
Published: (2025)
THGS: Lifelike Talking Human Avatar Synthesis From Monocular Video Via 3D Gaussian Splatting
by: Chuang Chen, et al.
Published: (2025)
by: Chuang Chen, et al.
Published: (2025)
LassoNet: Deep Lasso-Selection of 3D Point Clouds
by: Zhu-Tian, Chen, et al.
Published: (2019)
by: Zhu-Tian, Chen, et al.
Published: (2019)
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
by: Kelly, Chris, et al.
Published: (2024)
by: Kelly, Chris, et al.
Published: (2024)
3D-HGS: 3D Half-Gaussian Splatting
by: Li, Haolin, et al.
Published: (2024)
by: Li, Haolin, et al.
Published: (2024)
MotionDreamer: Exploring Semantic Video Diffusion features for Zero-Shot 3D Mesh Animation
by: Uzolas, Lukas, et al.
Published: (2024)
by: Uzolas, Lukas, et al.
Published: (2024)
Neural 3D Strokes: Creating Stylized 3D Scenes with Vectorized 3D Strokes
by: Duan, Hao-Bin, et al.
Published: (2023)
by: Duan, Hao-Bin, et al.
Published: (2023)
Img2CAD: Conditioned 3D CAD Model Generation from Single Image with Structured Visual Geometry
by: Chen, Tianrun, et al.
Published: (2024)
by: Chen, Tianrun, et al.
Published: (2024)
3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds
by: Sun, Fan-Yun, et al.
Published: (2025)
by: Sun, Fan-Yun, et al.
Published: (2025)
C3Editor: Achieving Controllable Consistency in 2D Model for 3D Editing
by: Tao, Zeng, et al.
Published: (2025)
by: Tao, Zeng, et al.
Published: (2025)
WonderHuman: Hallucinating Unseen Parts in Dynamic 3D Human Reconstruction
by: Wang, Zilong, et al.
Published: (2025)
by: Wang, Zilong, et al.
Published: (2025)
PointLLM-R: Enhancing 3D Point Cloud Reasoning via Chain-of-Thought
by: Chen, Chaoqi, et al.
Published: (2026)
by: Chen, Chaoqi, et al.
Published: (2026)
Matrix-3D: Omnidirectional Explorable 3D World Generation
by: Yang, Zhongqi, et al.
Published: (2025)
by: Yang, Zhongqi, et al.
Published: (2025)
Vision6D: 3D-to-2D Interactive Visualization and Annotation Tool for 6D Pose Estimation
by: Zhang, Yike, et al.
Published: (2025)
by: Zhang, Yike, et al.
Published: (2025)
Enforcing View-Consistency in Class-Agnostic 3D Segmentation Fields
by: Dumery, Corentin, et al.
Published: (2024)
by: Dumery, Corentin, et al.
Published: (2024)
RAVEN: Robust Advertisement Video Violation Temporal Grounding via Reinforcement Reasoning
by: Ji, Deyi, et al.
Published: (2025)
by: Ji, Deyi, et al.
Published: (2025)
SplatMesh: Interactive 3D Segmentation and Editing Using Mesh-Based Gaussian Splatting
by: Zhou, Kaichen, et al.
Published: (2023)
by: Zhou, Kaichen, et al.
Published: (2023)
Seamless and Aligned Texture Optimization for 3D Reconstruction
by: Lei Wang, et al.
Published: (2024)
by: Lei Wang, et al.
Published: (2024)
PartMotionEdit: Fine-Grained Text-Driven 3D Human Motion Editing via Part-Level Modulation
by: Yang, Yujie, et al.
Published: (2025)
by: Yang, Yujie, et al.
Published: (2025)
Causal Reasoning Elicits Controllable 3D Scene Generation
by: Chen, Shen, et al.
Published: (2025)
by: Chen, Shen, et al.
Published: (2025)
Similar Items
-
xLSTM-UNet can be an Effective 2D & 3D Medical Image Segmentation Backbone with Vision-LSTM (ViL) better than its Mamba Counterpart
by: Chen, Tianrun, et al.
Published: (2024) -
Adversarial Unsupervised Domain Adaptation for 3D Semantic Segmentation with 2D Image Fusion of Dense Depth
by: Xindan Zhang, et al.
Published: (2024) -
SAM2-Adapter: Evaluating & Adapting Segment Anything 2 in Downstream Tasks: Camouflage, Shadow, Medical Image Segmentation, and More
by: Chen, Tianrun, et al.
Published: (2024) -
Taming Video Models for 3D and 4D Generation via Zero-Shot Camera Control
by: Song, Chenxi, et al.
Published: (2025) -
AdaToken-3D: Dynamic Spatial Gating for Efficient 3D Large Multimodal-Models Reasoning
by: Zhang, Kai, et al.
Published: (2025)