ASHiTA: Automatic Scene-grounded HIerarchical Task Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Chang, Yun, Fermoselle, Leonor, Ta, Duy, Bucher, Bernadette, Carlone, Luca, Wang, Jiuguang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Zero-shot Object-Centric Instruction Following: Integrating Foundation Models with Traditional Navigation
by: Raychaudhuri, Sonia, et al.
Published: (2024)
by: Raychaudhuri, Sonia, et al.
Published: (2024)
CuriousBot: Interactive Mobile Exploration via Actionable 3D Relational Object Graph
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
Pandora: Articulated 3D Scene Graphs from Egocentric Vision
by: Yu, Alan, et al.
Published: (2026)
by: Yu, Alan, et al.
Published: (2026)
VGGT-SLAM 2.0: Real-time Dense Feed-forward Scene Reconstruction
by: Maggio, Dominic, et al.
Published: (2026)
by: Maggio, Dominic, et al.
Published: (2026)
Continuously Improving Mobile Manipulation with Autonomous Real-World RL
by: Mendonca, Russell, et al.
Published: (2024)
by: Mendonca, Russell, et al.
Published: (2024)
Towards Zero-Shot Point Cloud Registration Across Diverse Scales, Scenes, and Sensor Setups
by: Lim, Hyungtae, et al.
Published: (2026)
by: Lim, Hyungtae, et al.
Published: (2026)
CUPS: Improving Human Pose-Shape Estimators with Conformalized Deep Uncertainty
by: Zhang, Harry, et al.
Published: (2024)
by: Zhang, Harry, et al.
Published: (2024)
Uncertainty Quantification for Visual Object Pose Estimation
by: Shaikewitz, Lorenzo, et al.
Published: (2025)
by: Shaikewitz, Lorenzo, et al.
Published: (2025)
Category-Level Object Shape and Pose Estimation in Less Than a Millisecond
by: Shaikewitz, Lorenzo, et al.
Published: (2025)
by: Shaikewitz, Lorenzo, et al.
Published: (2025)
Picasso: Holistic Scene Reconstruction with Physics-Constrained Sampling
by: Yu, Xihang, et al.
Published: (2026)
by: Yu, Xihang, et al.
Published: (2026)
BUFFER-X: Towards Zero-Shot Point Cloud Registration in Diverse Scenes
by: Seo, Minkyun, et al.
Published: (2025)
by: Seo, Minkyun, et al.
Published: (2025)
Multi-Model 3D Registration: Finding Multiple Moving Objects in Cluttered Point Clouds
by: Jin, David, et al.
Published: (2024)
by: Jin, David, et al.
Published: (2024)
Describe Anything Anywhere At Any Moment
by: Gorlo, Nicolas, et al.
Published: (2025)
by: Gorlo, Nicolas, et al.
Published: (2025)
CRISP: Object Pose and Shape Estimation with Test-Time Adaptation
by: Shi, Jingnan, et al.
Published: (2024)
by: Shi, Jingnan, et al.
Published: (2024)
Test-Time Certifiable Self-Supervision to Bridge the Sim2Real Gap in Event-Based Satellite Pose Estimation
by: Jawaid, Mohsi, et al.
Published: (2024)
by: Jawaid, Mohsi, et al.
Published: (2024)
Efficient Multi-Task Scene Analysis with RGB-D Transformers
by: Fischedick, Söhnke Benedikt, et al.
Published: (2023)
by: Fischedick, Söhnke Benedikt, et al.
Published: (2023)
KISS-Matcher: Fast and Robust Point Cloud Registration Revisited
by: Lim, Hyungtae, et al.
Published: (2024)
by: Lim, Hyungtae, et al.
Published: (2024)
Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
by: Yang, Yurou, et al.
Published: (2026)
by: Yang, Yurou, et al.
Published: (2026)
DrivingScene: A Multi-Task Online Feed-Forward 3D Gaussian Splatting Method for Dynamic Driving Scenes
by: Hou, Qirui, et al.
Published: (2025)
by: Hou, Qirui, et al.
Published: (2025)
Bayesian Fields: Task-driven Open-Set Semantic Gaussian Splatting
by: Maggio, Dominic, et al.
Published: (2025)
by: Maggio, Dominic, et al.
Published: (2025)
GenDP: 3D Semantic Fields for Category-Level Generalizable Diffusion Policy
by: Wang, Yixuan, et al.
Published: (2024)
by: Wang, Yixuan, et al.
Published: (2024)
MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning
by: Hao, Jinkun, et al.
Published: (2025)
by: Hao, Jinkun, et al.
Published: (2025)
Box Pose and Shape Estimation and Domain Adaptation for Large-Scale Warehouse Automation
by: Yu, Xihang, et al.
Published: (2025)
by: Yu, Xihang, et al.
Published: (2025)
Planning-Guided Diffusion Policy Learning for Generalizable Contact-Rich Bimanual Manipulation
by: Li, Xuanlin, et al.
Published: (2024)
by: Li, Xuanlin, et al.
Published: (2024)
Clutt3R-Seg: Sparse-view 3D Instance Segmentation for Language-grounded Grasping in Cluttered Scenes
by: Noh, Jeongho, et al.
Published: (2026)
by: Noh, Jeongho, et al.
Published: (2026)
SALT: A Flexible Semi-Automatic Labeling Tool for General LiDAR Point Clouds with Cross-Scene Adaptability and 4D Consistency
by: Wang, Yanbo, et al.
Published: (2025)
by: Wang, Yanbo, et al.
Published: (2025)
Mimicking-Bench: A Benchmark for Generalizable Humanoid-Scene Interaction Learning via Human Mimicking
by: Liu, Yun, et al.
Published: (2024)
by: Liu, Yun, et al.
Published: (2024)
TaskGround: Structured Executable Task Inference for Full-Scene Household Reasoning
by: Feng, ZhiYuan, et al.
Published: (2026)
by: Feng, ZhiYuan, et al.
Published: (2026)
MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning
by: Ju, Yuanchen, et al.
Published: (2025)
by: Ju, Yuanchen, et al.
Published: (2025)
Motion Blender Gaussian Splatting for Dynamic Scene Reconstruction
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
Visually-grounded Humanoid Agents
by: Ye, Hang, et al.
Published: (2026)
by: Ye, Hang, et al.
Published: (2026)
Incremental Joint Learning of Depth, Pose and Implicit Scene Representation on Monocular Camera in Large-scale Scenes
by: Deng, Tianchen, et al.
Published: (2024)
by: Deng, Tianchen, et al.
Published: (2024)
RoadFormer: Duplex Transformer for RGB-Normal Semantic Road Scene Parsing
by: Li, Jiahang, et al.
Published: (2023)
by: Li, Jiahang, et al.
Published: (2023)
NovaFlow: Zero-Shot Manipulation via Actionable Flow from Generated Videos
by: Li, Hongyu, et al.
Published: (2025)
by: Li, Hongyu, et al.
Published: (2025)
InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts
by: Zhong, Weipeng, et al.
Published: (2025)
by: Zhong, Weipeng, et al.
Published: (2025)
Mono-Hydra++: Real-Time Monocular Scene Graph Construction with Multi-Task Learning for 3D Indoor Mapping
by: Udugama, U. V. B. L., et al.
Published: (2026)
by: Udugama, U. V. B. L., et al.
Published: (2026)
Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene Description
by: Halacheva, Anna-Maria, et al.
Published: (2024)
by: Halacheva, Anna-Maria, et al.
Published: (2024)
Estimating Commonsense Scene Composition on Belief Scene Graphs
by: Saucedo, Mario A. V., et al.
Published: (2025)
by: Saucedo, Mario A. V., et al.
Published: (2025)
WaterScenes: A Multi-Task 4D Radar-Camera Fusion Dataset and Benchmarks for Autonomous Driving on Water Surfaces
by: Yao, Shanliang, et al.
Published: (2023)
by: Yao, Shanliang, et al.
Published: (2023)
HIVE: HIerarchical Volume Encoding for Neural Implicit Surface Reconstruction
by: Gu, Xiaodong, et al.
Published: (2024)
by: Gu, Xiaodong, et al.
Published: (2024)
Similar Items
-
Zero-shot Object-Centric Instruction Following: Integrating Foundation Models with Traditional Navigation
by: Raychaudhuri, Sonia, et al.
Published: (2024) -
CuriousBot: Interactive Mobile Exploration via Actionable 3D Relational Object Graph
by: Wang, Yixuan, et al.
Published: (2025) -
Pandora: Articulated 3D Scene Graphs from Egocentric Vision
by: Yu, Alan, et al.
Published: (2026) -
VGGT-SLAM 2.0: Real-time Dense Feed-forward Scene Reconstruction
by: Maggio, Dominic, et al.
Published: (2026) -
Continuously Improving Mobile Manipulation with Autonomous Real-World RL
by: Mendonca, Russell, et al.
Published: (2024)