pySpatial: Generating 3D Visual Programs for Zero-Shot Spatial Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luo, Zhanpeng, Zhang, Ce, Yong, Silong, Dai, Cunxi, Wang, Qianwei, Ran, Haoxi, Shi, Guanya, Sycara, Katia, Xie, Yaqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GL-NeRF: Gauss-Laguerre Quadrature Enables Training-Free NeRF Acceleration
von: Yong, Silong, et al.
Veröffentlicht: (2024)
von: Yong, Silong, et al.
Veröffentlicht: (2024)
Enhancing Vision-Language Few-Shot Adaptation with Negative Learning
von: Zhang, Ce, et al.
Veröffentlicht: (2024)
von: Zhang, Ce, et al.
Veröffentlicht: (2024)
Instant4D: 4D Gaussian Splatting in Minutes
von: Luo, Zhanpeng, et al.
Veröffentlicht: (2025)
von: Luo, Zhanpeng, et al.
Veröffentlicht: (2025)
Dual Prototype Evolving for Test-Time Generalization of Vision-Language Models
von: Zhang, Ce, et al.
Veröffentlicht: (2024)
von: Zhang, Ce, et al.
Veröffentlicht: (2024)
Unifying Deep Predicate Invention with Pre-trained Foundation Models
von: Wang, Qianwei, et al.
Veröffentlicht: (2025)
von: Wang, Qianwei, et al.
Veröffentlicht: (2025)
Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation
von: Wan, Zifu, et al.
Veröffentlicht: (2024)
von: Wan, Zifu, et al.
Veröffentlicht: (2024)
Spectral-Aware Global Fusion for RGB-Thermal Semantic Segmentation
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
HiKER-SGG: Hierarchical Knowledge Enhanced Robust Scene Graph Generation
von: Zhang, Ce, et al.
Veröffentlicht: (2024)
von: Zhang, Ce, et al.
Veröffentlicht: (2024)
Evolving Contextual Safety in Multi-Modal Large Language Models via Inference-Time Self-Reflective Memory
von: Zhang, Ce, et al.
Veröffentlicht: (2026)
von: Zhang, Ce, et al.
Veröffentlicht: (2026)
Aug3D: Augmenting large scale outdoor datasets for Generalizable Novel View Synthesis
von: Rauniyar, Aditya, et al.
Veröffentlicht: (2025)
von: Rauniyar, Aditya, et al.
Veröffentlicht: (2025)
Jailbreaking Frontier Foundation Models Through Intention Deception
von: Wang, Xinhe, et al.
Veröffentlicht: (2026)
von: Wang, Xinhe, et al.
Veröffentlicht: (2026)
ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models
von: Wan, Zifu, et al.
Veröffentlicht: (2025)
von: Wan, Zifu, et al.
Veröffentlicht: (2025)
OMG: Opacity Matters in Material Modeling with Gaussian Splatting
von: Yong, Silong, et al.
Veröffentlicht: (2025)
von: Yong, Silong, et al.
Veröffentlicht: (2025)
ShapeGrasp: Zero-Shot Task-Oriented Grasping with Large Language Models through Geometric Decomposition
von: Li, Samuel, et al.
Veröffentlicht: (2024)
von: Li, Samuel, et al.
Veröffentlicht: (2024)
InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning
von: Wan, Zifu, et al.
Veröffentlicht: (2025)
von: Wan, Zifu, et al.
Veröffentlicht: (2025)
Modeling Latent Partner Strategies for Adaptive Zero-Shot Human-Agent Collaboration
von: Li, Benjamin, et al.
Veröffentlicht: (2025)
von: Li, Benjamin, et al.
Veröffentlicht: (2025)
Generalizable Dense Reward for Long-Horizon Robotic Tasks
von: Yong, Silong, et al.
Veröffentlicht: (2026)
von: Yong, Silong, et al.
Veröffentlicht: (2026)
Theory of Mind Guided Strategy Adaptation for Zero-Shot Coordination
von: Ni, Andrew, et al.
Veröffentlicht: (2026)
von: Ni, Andrew, et al.
Veröffentlicht: (2026)
SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning
von: Chunhachatrachai, Pawat, et al.
Veröffentlicht: (2026)
von: Chunhachatrachai, Pawat, et al.
Veröffentlicht: (2026)
SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual Grounding
von: Jin, Zhao, et al.
Veröffentlicht: (2025)
von: Jin, Zhao, et al.
Veröffentlicht: (2025)
B3C: A Minimalist Approach to Offline Multi-Agent Reinforcement Learning
von: Kim, Woojun, et al.
Veröffentlicht: (2025)
von: Kim, Woojun, et al.
Veröffentlicht: (2025)
Fair Cooperation in Mixed-Motive Games via Conflict-Aware Gradient Adjustment
von: Kim, Woojun, et al.
Veröffentlicht: (2025)
von: Kim, Woojun, et al.
Veröffentlicht: (2025)
Spatial-VLN: Zero-Shot Vision-and-Language Navigation With Explicit Spatial Perception and Exploration
von: Yue, Lu, et al.
Veröffentlicht: (2026)
von: Yue, Lu, et al.
Veröffentlicht: (2026)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
von: Zhang, Jiwen, et al.
Veröffentlicht: (2026)
von: Zhang, Jiwen, et al.
Veröffentlicht: (2026)
Improved Visual-Spatial Reasoning via R1-Zero-Like Training
von: Liao, Zhenyi, et al.
Veröffentlicht: (2025)
von: Liao, Zhenyi, et al.
Veröffentlicht: (2025)
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models
von: Taguchi, Shun, et al.
Veröffentlicht: (2025)
von: Taguchi, Shun, et al.
Veröffentlicht: (2025)
SEE-2-SOUND: Zero-Shot Spatial Environment-to-Spatial Sound
von: Dagli, Rishit, et al.
Veröffentlicht: (2024)
von: Dagli, Rishit, et al.
Veröffentlicht: (2024)
GSMem: 3D Gaussian Splatting as Persistent Spatial Memory for Zero-Shot Embodied Exploration and Reasoning
von: Lu, Yiren, et al.
Veröffentlicht: (2026)
von: Lu, Yiren, et al.
Veröffentlicht: (2026)
SpatialAnt: Autonomous Zero-Shot Robot Navigation via Active Scene Reconstruction and Visual Anticipation
von: Zhang, Jiwen, et al.
Veröffentlicht: (2026)
von: Zhang, Jiwen, et al.
Veröffentlicht: (2026)
HiMemFormer: Hierarchical Memory-Aware Transformer for Multi-Agent Action Anticipation
von: Wang, Zirui, et al.
Veröffentlicht: (2024)
von: Wang, Zirui, et al.
Veröffentlicht: (2024)
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
von: Ma, Wufei, et al.
Veröffentlicht: (2025)
von: Ma, Wufei, et al.
Veröffentlicht: (2025)
MSNav: Zero-Shot Vision-and-Language Navigation with Dynamic Memory and LLM Spatial Reasoning
von: Liu, Chenghao, et al.
Veröffentlicht: (2025)
von: Liu, Chenghao, et al.
Veröffentlicht: (2025)
ReplicateAnyScene: Zero-Shot Video-to-3D Composition via Textual-Visual-Spatial Alignment
von: Dong, Mingyu, et al.
Veröffentlicht: (2026)
von: Dong, Mingyu, et al.
Veröffentlicht: (2026)
Direct Numerical Layout Generation for 3D Indoor Scene Synthesis via Spatial Reasoning
von: Ran, Xingjian, et al.
Veröffentlicht: (2025)
von: Ran, Xingjian, et al.
Veröffentlicht: (2025)
BFM-Zero: A Promptable Behavioral Foundation Model for Humanoid Control Using Unsupervised Reinforcement Learning
von: Li, Yitang, et al.
Veröffentlicht: (2025)
von: Li, Yitang, et al.
Veröffentlicht: (2025)
Multi-Robot Navigation in Social Mini-Games: Definitions, Taxonomy, and Algorithms
von: Chandra, Rohan, et al.
Veröffentlicht: (2025)
von: Chandra, Rohan, et al.
Veröffentlicht: (2025)
Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
Reinforcement Learning with Physics-Informed Symbolic Program Priors for Zero-Shot Wireless Indoor Navigation
von: Li, Tao, et al.
Veröffentlicht: (2025)
von: Li, Tao, et al.
Veröffentlicht: (2025)
Open-World Visual Reasoning by a Neuro-Symbolic Program of Zero-Shot Symbols
von: Burghouts, Gertjan, et al.
Veröffentlicht: (2024)
von: Burghouts, Gertjan, et al.
Veröffentlicht: (2024)
Reconfigurable Robot Control Using Flexible Coupling Mechanisms
von: Yi, Sha, et al.
Veröffentlicht: (2023)
von: Yi, Sha, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
GL-NeRF: Gauss-Laguerre Quadrature Enables Training-Free NeRF Acceleration
von: Yong, Silong, et al.
Veröffentlicht: (2024) -
Enhancing Vision-Language Few-Shot Adaptation with Negative Learning
von: Zhang, Ce, et al.
Veröffentlicht: (2024) -
Instant4D: 4D Gaussian Splatting in Minutes
von: Luo, Zhanpeng, et al.
Veröffentlicht: (2025) -
Dual Prototype Evolving for Test-Time Generalization of Vision-Language Models
von: Zhang, Ce, et al.
Veröffentlicht: (2024) -
Unifying Deep Predicate Invention with Pre-trained Foundation Models
von: Wang, Qianwei, et al.
Veröffentlicht: (2025)