What Do Latent Action Models Actually Learn?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Chuheng, Pearce, Tim, Zhang, Pushi, Wang, Kaixin, Chen, Xiaoyu, Shen, Wei, Zhao, Li, Bian, Jiang
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911260296085504
author Zhang, Chuheng
Pearce, Tim
Zhang, Pushi
Wang, Kaixin
Chen, Xiaoyu
Shen, Wei
Zhao, Li
Bian, Jiang
author_facet Zhang, Chuheng
Pearce, Tim
Zhang, Pushi
Wang, Kaixin
Chen, Xiaoyu
Shen, Wei
Zhao, Li
Bian, Jiang
contents Latent action models (LAMs) aim to learn action-relevant changes from unlabeled videos by compressing changes between frames as latents. However, differences between video frames can be caused by controllable changes as well as exogenous noise, leading to an important concern -- do latents capture the changes caused by actions or irrelevant noise? This paper studies this issue analytically, presenting a linear model that encapsulates the essence of LAM learning, while being tractable.This provides several insights, including connections between LAM and principal component analysis (PCA), desiderata of the data-generating policy, and justification of strategies to encourage learning controllable changes using data augmentation, data cleaning, and auxiliary action-prediction. We also provide illustrative results based on numerical simulation, shedding light on the specific structure of observations, actions, and noise in data that influence LAM learning.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15691
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle What Do Latent Action Models Actually Learn?
Zhang, Chuheng
Pearce, Tim
Zhang, Pushi
Wang, Kaixin
Chen, Xiaoyu
Shen, Wei
Zhao, Li
Bian, Jiang
Machine Learning
Artificial Intelligence
Latent action models (LAMs) aim to learn action-relevant changes from unlabeled videos by compressing changes between frames as latents. However, differences between video frames can be caused by controllable changes as well as exogenous noise, leading to an important concern -- do latents capture the changes caused by actions or irrelevant noise? This paper studies this issue analytically, presenting a linear model that encapsulates the essence of LAM learning, while being tractable.This provides several insights, including connections between LAM and principal component analysis (PCA), desiderata of the data-generating policy, and justification of strategies to encourage learning controllable changes using data augmentation, data cleaning, and auxiliary action-prediction. We also provide illustrative results based on numerical simulation, shedding light on the specific structure of observations, actions, and noise in data that influence LAM learning.
title What Do Latent Action Models Actually Learn?
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.15691