3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Fan-Yun, Wu, Shengguang, Jacobsen, Christian, Yim, Thomas, Zou, Haoming, Zook, Alex, Li, Shangru, Chou, Yu-Hsin, Can, Ethem, Wu, Xunlei, Eppner, Clemens, Blukis, Valts, Tremblay, Jonathan, Wu, Jiajun, Birchfield, Stan, Haber, Nick |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GRS: Generating Robotic Simulation Tasks from Real-World Images
by: Zook, Alex, et al.
Published: (2024)
by: Zook, Alex, et al.
Published: (2024)
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
by: Song, Chan Hee, et al.
Published: (2024)
by: Song, Chan Hee, et al.
Published: (2024)
OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
by: Singh, Ishika, et al.
Published: (2025)
by: Singh, Ishika, et al.
Published: (2025)
Neural Implicit Representation for Building Digital Twins of Unknown Articulated Objects
by: Weng, Yijia, et al.
Published: (2024)
by: Weng, Yijia, et al.
Published: (2024)
Robot Policy Evaluation for Sim-to-Real Transfer: A Benchmarking Perspective
by: Yang, Xuning, et al.
Published: (2025)
by: Yang, Xuning, et al.
Published: (2025)
RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
by: Yang, Xuning, et al.
Published: (2026)
by: Yang, Xuning, et al.
Published: (2026)
Partial-View Object View Synthesis via Filtered Inversion
by: Sun, Fan-Yun, et al.
Published: (2023)
by: Sun, Fan-Yun, et al.
Published: (2023)
BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
by: Bhat, Vineet, et al.
Published: (2025)
by: Bhat, Vineet, et al.
Published: (2025)
3D-MVP: 3D Multiview Pretraining for Robotic Manipulation
by: Qian, Shengyi, et al.
Published: (2024)
by: Qian, Shengyi, et al.
Published: (2024)
Snap-it, Tap-it, Splat-it: Tactile-Informed 3D Gaussian Splatting for Reconstructing Challenging Surfaces
by: Comi, Mauro, et al.
Published: (2024)
by: Comi, Mauro, et al.
Published: (2024)
Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images
by: Wu, Shengguang, et al.
Published: (2025)
by: Wu, Shengguang, et al.
Published: (2025)
FactorSim: Generative Simulation via Factorized Representation
by: Sun, Fan-Yun, et al.
Published: (2024)
by: Sun, Fan-Yun, et al.
Published: (2024)
NeRFDeformer: NeRF Transformation from a Single View via 3D Scene Flows
by: Tang, Zhenggang, et al.
Published: (2024)
by: Tang, Zhenggang, et al.
Published: (2024)
GraspGen-X: Cross-Embodiment 6-DOF Diffusion-based Grasping
by: Han, Beining, et al.
Published: (2026)
by: Han, Beining, et al.
Published: (2026)
FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects
by: Wen, Bowen, et al.
Published: (2023)
by: Wen, Bowen, et al.
Published: (2023)
VLA-0: Building State-of-the-Art VLAs with Zero Modification
by: Goyal, Ankit, et al.
Published: (2025)
by: Goyal, Ankit, et al.
Published: (2025)
Fly, Fail, Fix: Iterative Game Repair with Reinforcement Learning and Large Multimodal Models
by: Zook, Alex, et al.
Published: (2025)
by: Zook, Alex, et al.
Published: (2025)
GraspGen: A Diffusion-based Framework for 6-DOF Grasping with On-Generator Training
by: Murali, Adithyavairavan, et al.
Published: (2025)
by: Murali, Adithyavairavan, et al.
Published: (2025)
LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models
by: Sun, Fan-Yun, et al.
Published: (2024)
by: Sun, Fan-Yun, et al.
Published: (2024)
One-Shot Transfer of Long-Horizon Extrinsic Manipulation Through Contact Retargeting
by: Wu, Albert, et al.
Published: (2024)
by: Wu, Albert, et al.
Published: (2024)
3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
by: Deng, Jiajun, et al.
Published: (2025)
by: Deng, Jiajun, et al.
Published: (2025)
cuRoboV2: Dynamics-Aware Motion Generation with Depth-Fused Distance Fields for High-DoF Robots
by: Sundaralingam, Balakumar, et al.
Published: (2026)
by: Sundaralingam, Balakumar, et al.
Published: (2026)
Fast-FoundationStereo: Real-Time Zero-Shot Stereo Matching
by: Wen, Bowen, et al.
Published: (2025)
by: Wen, Bowen, et al.
Published: (2025)
RVT-2: Learning Precise Manipulation from Few Demonstrations
by: Goyal, Ankit, et al.
Published: (2024)
by: Goyal, Ankit, et al.
Published: (2024)
Lifting Motion to the 3D World via 2D Diffusion
by: Li, Jiaman, et al.
Published: (2024)
by: Li, Jiaman, et al.
Published: (2024)
CT2Rep: Automated Radiology Report Generation for 3D Medical Imaging
by: Hamamci, Ibrahim Ethem, et al.
Published: (2024)
by: Hamamci, Ibrahim Ethem, et al.
Published: (2024)
Anymate: A Dataset and Baselines for Learning 3D Object Rigging
by: Deng, Yufan, et al.
Published: (2025)
by: Deng, Yufan, et al.
Published: (2025)
WonderZoom: Multi-Scale 3D World Generation
by: Cao, Jin, et al.
Published: (2025)
by: Cao, Jin, et al.
Published: (2025)
Arbitrary norm growth in the 3D Navier-Stokes equations
by: Palasek, Stan
Published: (2025)
by: Palasek, Stan
Published: (2025)
CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interaction
by: Xie, Xianghui, et al.
Published: (2025)
by: Xie, Xianghui, et al.
Published: (2025)
An Embodied Generalist Agent in 3D World
by: Huang, Jiangyong, et al.
Published: (2023)
by: Huang, Jiangyong, et al.
Published: (2023)
Real2SAM2Real: Generative 3D Caches as Complementary Context for Video Diffusion
by: Wu, Jiayi, et al.
Published: (2026)
by: Wu, Jiayi, et al.
Published: (2026)
Holodeck: Language Guided Generation of 3D Embodied AI Environments
by: Yang, Yue, et al.
Published: (2023)
by: Yang, Yue, et al.
Published: (2023)
Unveiling the Role of Lewis Base Strength in Small-Molecule Passivation of Defect Perovskites
by: Wu, Yi-Chen, et al.
Published: (2024)
by: Wu, Yi-Chen, et al.
Published: (2024)
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
by: Yuan, Wentao, et al.
Published: (2024)
by: Yuan, Wentao, et al.
Published: (2024)
Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography
by: Hamamci, Ibrahim Ethem, et al.
Published: (2024)
by: Hamamci, Ibrahim Ethem, et al.
Published: (2024)
RaySt3R: Predicting Novel Depth Maps for Zero-Shot Object Completion
by: Duisterhof, Bardienus P., et al.
Published: (2025)
by: Duisterhof, Bardienus P., et al.
Published: (2025)
Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners
by: Feng, Chun, et al.
Published: (2024)
by: Feng, Chun, et al.
Published: (2024)
Learning the 3D Fauna of the Web
by: Li, Zizhang, et al.
Published: (2024)
by: Li, Zizhang, et al.
Published: (2024)
Ponymation: Learning Articulated 3D Animal Motions from Unlabeled Online Videos
by: Sun, Keqiang, et al.
Published: (2023)
by: Sun, Keqiang, et al.
Published: (2023)
Similar Items
-
GRS: Generating Robotic Simulation Tasks from Real-World Images
by: Zook, Alex, et al.
Published: (2024) -
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
by: Song, Chan Hee, et al.
Published: (2024) -
OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
by: Singh, Ishika, et al.
Published: (2025) -
Neural Implicit Representation for Building Digital Twins of Unknown Articulated Objects
by: Weng, Yijia, et al.
Published: (2024) -
Robot Policy Evaluation for Sim-to-Real Transfer: A Benchmarking Perspective
by: Yang, Xuning, et al.
Published: (2025)