DecisionNCE: Embodied Multimodal Representations via Implicit Preference Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jianxiong, Zheng, Jinliang, Zheng, Yinan, Mao, Liyuan, Hu, Xiao, Cheng, Sijie, Niu, Haoyi, Liu, Jihao, Liu, Yu, Liu, Jingjing, Zhang, Ya-Qin, Zhan, Xianyuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910458480427008
author Li, Jianxiong
Zheng, Jinliang
Zheng, Yinan
Mao, Liyuan
Hu, Xiao
Cheng, Sijie
Niu, Haoyi
Liu, Jihao
Liu, Yu
Liu, Jingjing
Zhang, Ya-Qin
Zhan, Xianyuan
author_facet Li, Jianxiong
Zheng, Jinliang
Zheng, Yinan
Mao, Liyuan
Hu, Xiao
Cheng, Sijie
Niu, Haoyi
Liu, Jihao
Liu, Yu
Liu, Jingjing
Zhang, Ya-Qin
Zhan, Xianyuan
contents Multimodal pretraining is an effective strategy for the trinity of goals of representation learning in autonomous robots: 1) extracting both local and global task progressions; 2) enforcing temporal consistency of visual representation; 3) capturing trajectory-level language grounding. Most existing methods approach these via separate objectives, which often reach sub-optimal solutions. In this paper, we propose a universal unified objective that can simultaneously extract meaningful task progression information from image sequences and seamlessly align them with language instructions. We discover that via implicit preferences, where a visual trajectory inherently aligns better with its corresponding language instruction than mismatched pairs, the popular Bradley-Terry model can transform into representation learning through proper reward reparameterizations. The resulted framework, DecisionNCE, mirrors an InfoNCE-style objective but is distinctively tailored for decision-making tasks, providing an embodied representation learning framework that elegantly extracts both local and global task progression features, with temporal consistency enforced through implicit time contrastive learning, while ensuring trajectory-level instruction grounding via multimodal joint encoding. Evaluation on both simulated and real robots demonstrates that DecisionNCE effectively facilitates diverse downstream policy learning tasks, offering a versatile solution for unified representation and reward learning. Project Page: https://2toinf.github.io/DecisionNCE/
format Preprint
id arxiv_https___arxiv_org_abs_2402_18137
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DecisionNCE: Embodied Multimodal Representations via Implicit Preference Learning
Li, Jianxiong
Zheng, Jinliang
Zheng, Yinan
Mao, Liyuan
Hu, Xiao
Cheng, Sijie
Niu, Haoyi
Liu, Jihao
Liu, Yu
Liu, Jingjing
Zhang, Ya-Qin
Zhan, Xianyuan
Robotics
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
Multimodal pretraining is an effective strategy for the trinity of goals of representation learning in autonomous robots: 1) extracting both local and global task progressions; 2) enforcing temporal consistency of visual representation; 3) capturing trajectory-level language grounding. Most existing methods approach these via separate objectives, which often reach sub-optimal solutions. In this paper, we propose a universal unified objective that can simultaneously extract meaningful task progression information from image sequences and seamlessly align them with language instructions. We discover that via implicit preferences, where a visual trajectory inherently aligns better with its corresponding language instruction than mismatched pairs, the popular Bradley-Terry model can transform into representation learning through proper reward reparameterizations. The resulted framework, DecisionNCE, mirrors an InfoNCE-style objective but is distinctively tailored for decision-making tasks, providing an embodied representation learning framework that elegantly extracts both local and global task progression features, with temporal consistency enforced through implicit time contrastive learning, while ensuring trajectory-level instruction grounding via multimodal joint encoding. Evaluation on both simulated and real robots demonstrates that DecisionNCE effectively facilitates diverse downstream policy learning tasks, offering a versatile solution for unified representation and reward learning. Project Page: https://2toinf.github.io/DecisionNCE/
title DecisionNCE: Embodied Multimodal Representations via Implicit Preference Learning
topic Robotics
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2402.18137