Temporal Representation Alignment: Successor Features Enable Emergent Compositionality in Robot Instruction Following

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Myers, Vivek, Zheng, Bill Chunyuan, Dragan, Anca, Fang, Kuan, Levine, Sergey
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912230796165120
author Myers, Vivek
Zheng, Bill Chunyuan
Dragan, Anca
Fang, Kuan
Levine, Sergey
author_facet Myers, Vivek
Zheng, Bill Chunyuan
Dragan, Anca
Fang, Kuan
Levine, Sergey
contents Effective task representations should facilitate compositionality, such that after learning a variety of basic tasks, an agent can perform compound tasks consisting of multiple steps simply by composing the representations of the constituent steps together. While this is conceptually simple and appealing, it is not clear how to automatically learn representations that enable this sort of compositionality. We show that learning to associate the representations of current and future states with a temporal alignment loss can improve compositional generalization, even in the absence of any explicit subtask planning or reinforcement learning. We evaluate our approach across diverse robotic manipulation tasks as well as in simulation, showing substantial improvements for tasks specified with either language or goal images.
format Preprint
id arxiv_https___arxiv_org_abs_2502_05454
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Temporal Representation Alignment: Successor Features Enable Emergent Compositionality in Robot Instruction Following
Myers, Vivek
Zheng, Bill Chunyuan
Dragan, Anca
Fang, Kuan
Levine, Sergey
Robotics
Machine Learning
Effective task representations should facilitate compositionality, such that after learning a variety of basic tasks, an agent can perform compound tasks consisting of multiple steps simply by composing the representations of the constituent steps together. While this is conceptually simple and appealing, it is not clear how to automatically learn representations that enable this sort of compositionality. We show that learning to associate the representations of current and future states with a temporal alignment loss can improve compositional generalization, even in the absence of any explicit subtask planning or reinforcement learning. We evaluate our approach across diverse robotic manipulation tasks as well as in simulation, showing substantial improvements for tasks specified with either language or goal images.
title Temporal Representation Alignment: Successor Features Enable Emergent Compositionality in Robot Instruction Following
topic Robotics
Machine Learning
url https://arxiv.org/abs/2502.05454