Language Control Diffusion: Efficiently Scaling through Space, Time, and Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Edwin, Lu, Yujie, Huang, Shinda, Wang, William, Zhang, Amy
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915053705363456
author Zhang, Edwin
Lu, Yujie
Huang, Shinda
Wang, William
Zhang, Amy
author_facet Zhang, Edwin
Lu, Yujie
Huang, Shinda
Wang, William
Zhang, Amy
contents Training generalist agents is difficult across several axes, requiring us to deal with high-dimensional inputs (space), long horizons (time), and generalization to novel tasks. Recent advances with architectures have allowed for improved scaling along one or two of these axes, but are still computationally prohibitive to use. In this paper, we propose to address all three axes by leveraging \textbf{L}anguage to \textbf{C}ontrol \textbf{D}iffusion models as a hierarchical planner conditioned on language (LCD). We effectively and efficiently scale diffusion models for planning in extended temporal, state, and task dimensions to tackle long horizon control problems conditioned on natural language instructions, as a step towards generalist agents. Comparing LCD with other state-of-the-art models on the CALVIN language robotics benchmark finds that LCD outperforms other SOTA methods in multi-task success rates, whilst improving inference speed over other comparable diffusion models by 3.3x~15x. We show that LCD can successfully leverage the unique strength of diffusion models to produce coherent long range plans while addressing their weakness in generating low-level details and control.
format Preprint
id arxiv_https___arxiv_org_abs_2210_15629
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Language Control Diffusion: Efficiently Scaling through Space, Time, and Tasks
Zhang, Edwin
Lu, Yujie
Huang, Shinda
Wang, William
Zhang, Amy
Machine Learning
Artificial Intelligence
Computation and Language
Training generalist agents is difficult across several axes, requiring us to deal with high-dimensional inputs (space), long horizons (time), and generalization to novel tasks. Recent advances with architectures have allowed for improved scaling along one or two of these axes, but are still computationally prohibitive to use. In this paper, we propose to address all three axes by leveraging \textbf{L}anguage to \textbf{C}ontrol \textbf{D}iffusion models as a hierarchical planner conditioned on language (LCD). We effectively and efficiently scale diffusion models for planning in extended temporal, state, and task dimensions to tackle long horizon control problems conditioned on natural language instructions, as a step towards generalist agents. Comparing LCD with other state-of-the-art models on the CALVIN language robotics benchmark finds that LCD outperforms other SOTA methods in multi-task success rates, whilst improving inference speed over other comparable diffusion models by 3.3x~15x. We show that LCD can successfully leverage the unique strength of diffusion models to produce coherent long range plans while addressing their weakness in generating low-level details and control.
title Language Control Diffusion: Efficiently Scaling through Space, Time, and Tasks
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2210.15629