Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shen, Yutong, Liu, Hangxu, Zhang, Lei, Liu, Penghui, Xia, Ruizhe, Yao, Tianyi, Feng, Tongtong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2508.07842
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911167584141312
author Shen, Yutong
Liu, Hangxu
Zhang, Lei
Liu, Penghui
Xia, Ruizhe
Yao, Tianyi
Feng, Tongtong
author_facet Shen, Yutong
Liu, Hangxu
Zhang, Lei
Liu, Penghui
Xia, Ruizhe
Yao, Tianyi
Feng, Tongtong
contents Long-Horizon (LH) tasks in Human-Scene Interaction (HSI) are complex multi-step tasks that require continuous planning, sequential decision-making, and extended execution across domains to achieve the final goal. However, existing methods heavily rely on skill chaining by concatenating pre-trained subtasks, with environment observations and self-state tightly coupled, lacking the ability to generalize to new combinations of environments and skills, failing to complete various LH tasks across domains. To solve this problem, this paper presents DETACH, a cross-domain learning framework for LH tasks via biologically inspired dual-stream disentanglement. Inspired by the brain's "where-what" dual pathway mechanism, DETACH comprises two core modules: i) an environment learning module for spatial understanding, which captures object functions, spatial relationships, and scene semantics, achieving cross-domain transfer through complete environment-self disentanglement; ii) a skill learning module for task execution, which processes self-state information including joint degrees of freedom and motor patterns, enabling cross-skill transfer through independent motor pattern encoding. We conducted extensive experiments on various LH tasks in HSI scenes. Compared with existing methods, DETACH can achieve an average subtasks success rate improvement of 23% and average execution efficiency improvement of 29%.
format Preprint
id arxiv_https___arxiv_org_abs_2508_07842
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DETACH: Cross-domain Learning for Long-Horizon Tasks via Mixture of Disentangled Experts
Shen, Yutong
Liu, Hangxu
Zhang, Lei
Liu, Penghui
Xia, Ruizhe
Yao, Tianyi
Feng, Tongtong
Robotics
Artificial Intelligence
Long-Horizon (LH) tasks in Human-Scene Interaction (HSI) are complex multi-step tasks that require continuous planning, sequential decision-making, and extended execution across domains to achieve the final goal. However, existing methods heavily rely on skill chaining by concatenating pre-trained subtasks, with environment observations and self-state tightly coupled, lacking the ability to generalize to new combinations of environments and skills, failing to complete various LH tasks across domains. To solve this problem, this paper presents DETACH, a cross-domain learning framework for LH tasks via biologically inspired dual-stream disentanglement. Inspired by the brain's "where-what" dual pathway mechanism, DETACH comprises two core modules: i) an environment learning module for spatial understanding, which captures object functions, spatial relationships, and scene semantics, achieving cross-domain transfer through complete environment-self disentanglement; ii) a skill learning module for task execution, which processes self-state information including joint degrees of freedom and motor patterns, enabling cross-skill transfer through independent motor pattern encoding. We conducted extensive experiments on various LH tasks in HSI scenes. Compared with existing methods, DETACH can achieve an average subtasks success rate improvement of 23% and average execution efficiency improvement of 29%.
title DETACH: Cross-domain Learning for Long-Horizon Tasks via Mixture of Disentangled Experts
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2508.07842