TaskGround: Structured Executable Task Inference for Full-Scene Household Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, ZhiYuan, Deng, Yu, An, Ruichuan, Liu, Zhenhua, Li, Qixiu, Wu, Keming, Du, Zhiying, Wang, Weijie, Wang, Haoxiao, Chen, Shuang, Xu, Sicheng, Liang, Yaobo, Yang, Jiaolong, Guo, Baining |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
by: Feng, Zhiyuan, et al.
Published: (2025)
by: Feng, Zhiyuan, et al.
Published: (2025)
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
by: Shen, Yichao, et al.
Published: (2025)
by: Shen, Yichao, et al.
Published: (2025)
VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single Image
by: Xu, Sicheng, et al.
Published: (2025)
by: Xu, Sicheng, et al.
Published: (2025)
Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs
by: Xu, Sicheng, et al.
Published: (2026)
by: Xu, Sicheng, et al.
Published: (2026)
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
by: Li, Qixiu, et al.
Published: (2024)
by: Li, Qixiu, et al.
Published: (2024)
Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis
by: Zhang, Bowen, et al.
Published: (2025)
by: Zhang, Bowen, et al.
Published: (2025)
MobileManiBench: Simplifying Model Verification for Mobile Manipulation
by: Wang, Wenbo, et al.
Published: (2026)
by: Wang, Wenbo, et al.
Published: (2026)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
by: Li, Qixiu, et al.
Published: (2025)
by: Li, Qixiu, et al.
Published: (2025)
HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models
by: Liang, Huizhi, et al.
Published: (2026)
by: Liang, Huizhi, et al.
Published: (2026)
Cluster-Based Multi-Agent Task Scheduling for Space-Air-Ground Integrated Networks
by: Wang, Zhiying, et al.
Published: (2024)
by: Wang, Zhiying, et al.
Published: (2024)
Domain-Conditioned Scene Graphs for State-Grounded Task Planning
by: Herzog, Jonas, et al.
Published: (2025)
by: Herzog, Jonas, et al.
Published: (2025)
Anticipate & Act : Integrating LLMs and Classical Planning for Efficient Task Execution in Household Environments
by: Arora, Raghav, et al.
Published: (2025)
by: Arora, Raghav, et al.
Published: (2025)
Task-oriented Sequential Grounding and Navigation in 3D Scenes
by: Zhang, Zhuofan, et al.
Published: (2024)
by: Zhang, Zhuofan, et al.
Published: (2024)
When Robots Do the Chores: A Benchmark and Agent for Long-Horizon Household Task Execution
by: Zhu, Zilin, et al.
Published: (2026)
by: Zhu, Zilin, et al.
Published: (2026)
SWEET: Sparse World Modeling with Image Editing for Embodied Task Execution
by: Song, Yiren, et al.
Published: (2026)
by: Song, Yiren, et al.
Published: (2026)
SAGE: Scene Graph-Aware Guidance and Execution for Long-Horizon Manipulation Tasks
by: Li, Jialiang, et al.
Published: (2025)
by: Li, Jialiang, et al.
Published: (2025)
UniGraspTransformer: Simplified Policy Distillation for Scalable Dexterous Robotic Grasping
by: Wang, Wenbo, et al.
Published: (2024)
by: Wang, Wenbo, et al.
Published: (2024)
TaPS: A Performance Evaluation Suite for Task-based Execution Frameworks
by: Pauloski, J. Gregory, et al.
Published: (2024)
by: Pauloski, J. Gregory, et al.
Published: (2024)
VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time
by: Xu, Sicheng, et al.
Published: (2024)
by: Xu, Sicheng, et al.
Published: (2024)
MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision
by: Wang, Ruicheng, et al.
Published: (2024)
by: Wang, Ruicheng, et al.
Published: (2024)
The GA4GH Task Execution API: Enabling Easy Multi Cloud Task Execution
by: Kanitz, Alexander, et al.
Published: (2024)
by: Kanitz, Alexander, et al.
Published: (2024)
How Does Diverse Interpretability of Textual Prompts Impact Medical Vision-Language Zero-Shot Tasks?
by: Wang, Sicheng, et al.
Published: (2024)
by: Wang, Sicheng, et al.
Published: (2024)
LightPlanner: Unleashing the Reasoning Capabilities of Lightweight Large Language Models in Task Planning
by: Zhou, Weijie, et al.
Published: (2025)
by: Zhou, Weijie, et al.
Published: (2025)
HELP: Hierarchical Embodied Language Planner for Household Tasks
by: Korchemnyi, Alexandr V., et al.
Published: (2025)
by: Korchemnyi, Alexandr V., et al.
Published: (2025)
MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning
by: Hao, Jinkun, et al.
Published: (2025)
by: Hao, Jinkun, et al.
Published: (2025)
HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies
by: Du, Zhiying, et al.
Published: (2025)
by: Du, Zhiying, et al.
Published: (2025)
SEGT: A General Spatial Expansion Group Transformer for nuScenes Lidar-based Object Detection Task
by: Mei, Cheng, et al.
Published: (2024)
by: Mei, Cheng, et al.
Published: (2024)
Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA
by: Safwan, Itbaan, et al.
Published: (2025)
by: Safwan, Itbaan, et al.
Published: (2025)
CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
by: Dong, Qihua, et al.
Published: (2025)
by: Dong, Qihua, et al.
Published: (2025)
FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
by: Ju, Xuan, et al.
Published: (2025)
by: Ju, Xuan, et al.
Published: (2025)
Data-Locality-Aware Task Assignment and Scheduling for Distributed Job Executions
by: Zhao, Hailiang, et al.
Published: (2024)
by: Zhao, Hailiang, et al.
Published: (2024)
DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation
by: Wang, Yibo, et al.
Published: (2026)
by: Wang, Yibo, et al.
Published: (2026)
Hybrid and Oriented Harmonic Potentials for Safe Task Execution in Unknown Environment
by: Wang, Shuaikang, et al.
Published: (2023)
by: Wang, Shuaikang, et al.
Published: (2023)
ConceptAgent: LLM-Driven Precondition Grounding and Tree Search for Robust Task Planning and Execution
by: Rivera, Corban, et al.
Published: (2024)
by: Rivera, Corban, et al.
Published: (2024)
Task and Motion Planning for Execution in the Real
by: Pan, Tianyang, et al.
Published: (2024)
by: Pan, Tianyang, et al.
Published: (2024)
Relationship-Aware Hierarchical 3D Scene Graph for Task Reasoning
by: Puigjaner, Albert Gassol, et al.
Published: (2026)
by: Puigjaner, Albert Gassol, et al.
Published: (2026)
Disentangling Language Roles in Multilingual LLM Task Execution
by: Zhan, Qishi, et al.
Published: (2026)
by: Zhan, Qishi, et al.
Published: (2026)
Grounding Language Models in Autonomous Loco-manipulation Tasks
by: Wang, Jin, et al.
Published: (2024)
by: Wang, Jin, et al.
Published: (2024)
PPTC-R benchmark: Towards Evaluating the Robustness of Large Language Models for PowerPoint Task Completion
by: Zhang, Zekai, et al.
Published: (2024)
by: Zhang, Zekai, et al.
Published: (2024)
Beyond Task Diversity: Provable Representation Transfer for Sequential Multi-Task Linear Bandits
by: Duong, Thang, et al.
Published: (2025)
by: Duong, Thang, et al.
Published: (2025)
Similar Items
-
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
by: Feng, Zhiyuan, et al.
Published: (2025) -
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
by: Shen, Yichao, et al.
Published: (2025) -
VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single Image
by: Xu, Sicheng, et al.
Published: (2025) -
Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs
by: Xu, Sicheng, et al.
Published: (2026) -
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
by: Li, Qixiu, et al.
Published: (2024)