Planning with the Views via Scene Self-Exploration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Kangrui, Li, Linjie, Yang, Zhengyuan, Chen, Shiqi, Wang, Zihan, Fei-Fei, Li, Wu, Jiajun, Guibas, Leonidas, Wang, Lijuan, Li, Manling |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs
von: Hong, Yining, et al.
Veröffentlicht: (2026)
von: Hong, Yining, et al.
Veröffentlicht: (2026)
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
von: Hong, Yining, et al.
Veröffentlicht: (2026)
von: Hong, Yining, et al.
Veröffentlicht: (2026)
SparseDFF: Sparse-View Feature Distillation for One-Shot Dexterous Manipulation
von: Wang, Qianxu, et al.
Veröffentlicht: (2023)
von: Wang, Qianxu, et al.
Veröffentlicht: (2023)
Neural Attention Field: Emerging Point Relevance in 3D Scenes for One-Shot Dexterous Grasping
von: Wang, Qianxu, et al.
Veröffentlicht: (2024)
von: Wang, Qianxu, et al.
Veröffentlicht: (2024)
Beyond Words: Advancing Long-Text Image Generation via Multimodal Autoregressive Models
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2025)
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2025)
StereoPolicy: Improving Robotic Manipulation Policies via Stereo Perception
von: Han, Evans, et al.
Veröffentlicht: (2026)
von: Han, Evans, et al.
Veröffentlicht: (2026)
SAGE: Bridging Semantic and Actionable Parts for GEneralizable Manipulation of Articulated Objects
von: Geng, Haoran, et al.
Veröffentlicht: (2023)
von: Geng, Haoran, et al.
Veröffentlicht: (2023)
PhysPart: Physically Plausible Part Completion for Interactable Objects
von: Luo, Rundong, et al.
Veröffentlicht: (2024)
von: Luo, Rundong, et al.
Veröffentlicht: (2024)
MindCube: Spatial Mental Modeling from Limited Views
von: Wang, Qineng, et al.
Veröffentlicht: (2025)
von: Wang, Qineng, et al.
Veröffentlicht: (2025)
Hallucinating 360°: Panoramic Street-View Generation via Local Scenes Diffusion and Probabilistic Prompting
von: Teng, Fei, et al.
Veröffentlicht: (2025)
von: Teng, Fei, et al.
Veröffentlicht: (2025)
ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction
von: Wang, Qineng, et al.
Veröffentlicht: (2025)
von: Wang, Qineng, et al.
Veröffentlicht: (2025)
RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation
von: Kuang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Kuang, Yuxuan, et al.
Veröffentlicht: (2024)
D$^3$Fields: Dynamic 3D Descriptor Fields for Zero-Shot Generalizable Rearrangement
von: Wang, Yixuan, et al.
Veröffentlicht: (2023)
von: Wang, Yixuan, et al.
Veröffentlicht: (2023)
Lookahead Exploration with Neural Radiance Representation for Continuous Vision-Language Navigation
von: Wang, Zihan, et al.
Veröffentlicht: (2024)
von: Wang, Zihan, et al.
Veröffentlicht: (2024)
UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation
von: Tang, Yihe, et al.
Veröffentlicht: (2025)
von: Tang, Yihe, et al.
Veröffentlicht: (2025)
Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow
von: Dharmarajan, Karthik, et al.
Veröffentlicht: (2025)
von: Dharmarajan, Karthik, et al.
Veröffentlicht: (2025)
BFA: Best-Feature-Aware Fusion for Multi-View Fine-grained Manipulation
von: Lan, Zihan, et al.
Veröffentlicht: (2025)
von: Lan, Zihan, et al.
Veröffentlicht: (2025)
Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
von: Bai, Yongjie, et al.
Veröffentlicht: (2025)
von: Bai, Yongjie, et al.
Veröffentlicht: (2025)
Rodrigues Network for Learning Robot Actions
von: Zhang, Jialiang, et al.
Veröffentlicht: (2025)
von: Zhang, Jialiang, et al.
Veröffentlicht: (2025)
RoboEXP: Action-Conditioned Scene Graph via Interactive Exploration for Robotic Manipulation
von: Jiang, Hanxiao, et al.
Veröffentlicht: (2024)
von: Jiang, Hanxiao, et al.
Veröffentlicht: (2024)
Sim-to-Real Transfer via 3D Feature Fields for Vision-and-Language Navigation
von: Wang, Zihan, et al.
Veröffentlicht: (2024)
von: Wang, Zihan, et al.
Veröffentlicht: (2024)
Mimicking-Bench: A Benchmark for Generalizable Humanoid-Scene Interaction Learning via Human Mimicking
von: Liu, Yun, et al.
Veröffentlicht: (2024)
von: Liu, Yun, et al.
Veröffentlicht: (2024)
Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation
von: Yang, Zhengyuan, et al.
Veröffentlicht: (2023)
von: Yang, Zhengyuan, et al.
Veröffentlicht: (2023)
DOT-Sim: Differentiable Optical Tactile Simulation with Precise Real-to-Sim Physical Calibration
von: You, Yang, et al.
Veröffentlicht: (2026)
von: You, Yang, et al.
Veröffentlicht: (2026)
DVPE: Divided View Position Embedding for Multi-View 3D Object Detection
von: Wang, Jiasen, et al.
Veröffentlicht: (2024)
von: Wang, Jiasen, et al.
Veröffentlicht: (2024)
DiffPlace: Street View Generation via Place-Controllable Diffusion Model Enhancing Place Recognition
von: Li, Ji, et al.
Veröffentlicht: (2026)
von: Li, Ji, et al.
Veröffentlicht: (2026)
Controllable Pedestrian Video Editing for Multi-View Driving Scenarios via Motion Sequence
von: Fu, Danzhen, et al.
Veröffentlicht: (2025)
von: Fu, Danzhen, et al.
Veröffentlicht: (2025)
Boundary Exploration of Next Best View Policy in 3D Robotic Scanning
von: Li, Leihui, et al.
Veröffentlicht: (2024)
von: Li, Leihui, et al.
Veröffentlicht: (2024)
Monocular Semantic Scene Completion via Masked Recurrent Networks
von: Wang, Xuzhi, et al.
Veröffentlicht: (2025)
von: Wang, Xuzhi, et al.
Veröffentlicht: (2025)
Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos
von: Yuan, Chengbo, et al.
Veröffentlicht: (2024)
von: Yuan, Chengbo, et al.
Veröffentlicht: (2024)
Instance-aware Exploration-Verification-Exploitation for Instance ImageGoal Navigation
von: Lei, Xiaohan, et al.
Veröffentlicht: (2024)
von: Lei, Xiaohan, et al.
Veröffentlicht: (2024)
Not All Voxels Are Equal: Hardness-Aware Semantic Scene Completion with Self-Distillation
von: Wang, Song, et al.
Veröffentlicht: (2024)
von: Wang, Song, et al.
Veröffentlicht: (2024)
ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
von: Huang, Wenlong, et al.
Veröffentlicht: (2024)
von: Huang, Wenlong, et al.
Veröffentlicht: (2024)
Scene-Agnostic Traversability Labeling and Estimation via a Multimodal Self-supervised Framework
von: Fang, Zipeng, et al.
Veröffentlicht: (2025)
von: Fang, Zipeng, et al.
Veröffentlicht: (2025)
Bring Metric Functions into Diffusion Models
von: An, Jie, et al.
Veröffentlicht: (2024)
von: An, Jie, et al.
Veröffentlicht: (2024)
FastViDAR: Real-Time Omnidirectional Depth Estimation via Alternative Hierarchical Attention
von: Zhao, Hangtian, et al.
Veröffentlicht: (2025)
von: Zhao, Hangtian, et al.
Veröffentlicht: (2025)
SeFlow: A Self-Supervised Scene Flow Method in Autonomous Driving
von: Zhang, Qingwen, et al.
Veröffentlicht: (2024)
von: Zhang, Qingwen, et al.
Veröffentlicht: (2024)
GLEAM: Learning Generalizable Exploration Policy for Active Mapping in Complex 3D Indoor Scenes
von: Chen, Xiao, et al.
Veröffentlicht: (2025)
von: Chen, Xiao, et al.
Veröffentlicht: (2025)
MaMi-HOI: Harmonizing Global Kinematics and Local Geometry for Human-Object Interaction Generation
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
AERR-Nav: Adaptive Exploration-Recovery-Reminiscing Strategy for Zero-Shot Object Navigation
von: Huang, Jingzhi, et al.
Veröffentlicht: (2026)
von: Huang, Jingzhi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs
von: Hong, Yining, et al.
Veröffentlicht: (2026) -
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
von: Hong, Yining, et al.
Veröffentlicht: (2026) -
SparseDFF: Sparse-View Feature Distillation for One-Shot Dexterous Manipulation
von: Wang, Qianxu, et al.
Veröffentlicht: (2023) -
Neural Attention Field: Emerging Point Relevance in 3D Scenes for One-Shot Dexterous Grasping
von: Wang, Qianxu, et al.
Veröffentlicht: (2024) -
Beyond Words: Advancing Long-Text Image Generation via Multimodal Autoregressive Models
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2025)