Understanding Multimodal Failure in Action-Chunking Behavioral Cloning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mazza, Lorenzo, Datres, Massimiliano, Rodriguez, Ariel, Bodenstedt, Sebastian, Kutyniok, Gitta, Speidel, Stefanie
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910246135398400
author Mazza, Lorenzo
Datres, Massimiliano
Rodriguez, Ariel
Bodenstedt, Sebastian
Kutyniok, Gitta
Speidel, Stefanie
author_facet Mazza, Lorenzo
Datres, Massimiliano
Rodriguez, Ariel
Bodenstedt, Sebastian
Kutyniok, Gitta
Speidel, Stefanie
contents Behavioral cloning becomes difficult when the same observation admits several valid actions. We study this problem for action-chunking policies and show that different multimodal parameterizations fail in different ways. For latent-variable policies, posterior-prior regularization makes deployment-time sampling more reliable, but excessive regularization removes the action-conditioned information needed to distinguish demonstrated modes. Reducing this regularization can preserve mode information, but then success depends on whether the prior covers the relevant latent regions. For action-space generative policies, multimodality is constrained by the smoothness of the base-to-action transport: a map with small Lipschitz constant cannot assign substantial probability to many well-separated modes. Covering many modes therefore requires either sharp transitions in base space or off-support bridge regions in action space. Experiments on synthetic multimodal tasks and robotic simulation benchmarks support these mechanisms.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22493
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Understanding Multimodal Failure in Action-Chunking Behavioral Cloning
Mazza, Lorenzo
Datres, Massimiliano
Rodriguez, Ariel
Bodenstedt, Sebastian
Kutyniok, Gitta
Speidel, Stefanie
Machine Learning
Artificial Intelligence
Robotics
Behavioral cloning becomes difficult when the same observation admits several valid actions. We study this problem for action-chunking policies and show that different multimodal parameterizations fail in different ways. For latent-variable policies, posterior-prior regularization makes deployment-time sampling more reliable, but excessive regularization removes the action-conditioned information needed to distinguish demonstrated modes. Reducing this regularization can preserve mode information, but then success depends on whether the prior covers the relevant latent regions. For action-space generative policies, multimodality is constrained by the smoothness of the base-to-action transport: a map with small Lipschitz constant cannot assign substantial probability to many well-separated modes. Covering many modes therefore requires either sharp transitions in base space or off-support bridge regions in action space. Experiments on synthetic multimodal tasks and robotic simulation benchmarks support these mechanisms.
title Understanding Multimodal Failure in Action-Chunking Behavioral Cloning
topic Machine Learning
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2605.22493