State Combinatorial Generalization In Decision Making With Conditional Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duan, Xintong, He, Yutong, Tajwar, Fahim, Chen, Wen-Tse, Salakhutdinov, Ruslan, Schneider, Jeff
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908709854117888
author Duan, Xintong
He, Yutong
Tajwar, Fahim
Chen, Wen-Tse
Salakhutdinov, Ruslan
Schneider, Jeff
author_facet Duan, Xintong
He, Yutong
Tajwar, Fahim
Chen, Wen-Tse
Salakhutdinov, Ruslan
Schneider, Jeff
contents Many real-world decision-making problems are combinatorial in nature, where states (e.g., surrounding traffic of a self-driving car) can be seen as a combination of basic elements (e.g., pedestrians, trees, and other cars). Due to combinatorial complexity, observing all combinations of basic elements in the training set is infeasible, which leads to an essential yet understudied problem of zero-shot generalization to states that are unseen combinations of previously seen elements. In this work, we first formalize this problem and then demonstrate how existing value-based reinforcement learning (RL) algorithms struggle due to unreliable value predictions in unseen states. We argue that this problem cannot be addressed with exploration alone, but requires more expressive and generalizable models. We demonstrate that behavior cloning with a conditioned diffusion model trained on successful trajectory generalizes better to states formed by new combinations of seen elements than traditional RL methods. Through experiments in maze, driving, and multiagent environments, we show that conditioned diffusion models outperform traditional RL techniques and highlight the broad applicability of our problem formulation.
format Preprint
id arxiv_https___arxiv_org_abs_2501_13241
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle State Combinatorial Generalization In Decision Making With Conditional Diffusion Models
Duan, Xintong
He, Yutong
Tajwar, Fahim
Chen, Wen-Tse
Salakhutdinov, Ruslan
Schneider, Jeff
Machine Learning
Many real-world decision-making problems are combinatorial in nature, where states (e.g., surrounding traffic of a self-driving car) can be seen as a combination of basic elements (e.g., pedestrians, trees, and other cars). Due to combinatorial complexity, observing all combinations of basic elements in the training set is infeasible, which leads to an essential yet understudied problem of zero-shot generalization to states that are unseen combinations of previously seen elements. In this work, we first formalize this problem and then demonstrate how existing value-based reinforcement learning (RL) algorithms struggle due to unreliable value predictions in unseen states. We argue that this problem cannot be addressed with exploration alone, but requires more expressive and generalizable models. We demonstrate that behavior cloning with a conditioned diffusion model trained on successful trajectory generalizes better to states formed by new combinations of seen elements than traditional RL methods. Through experiments in maze, driving, and multiagent environments, we show that conditioned diffusion models outperform traditional RL techniques and highlight the broad applicability of our problem formulation.
title State Combinatorial Generalization In Decision Making With Conditional Diffusion Models
topic Machine Learning
url https://arxiv.org/abs/2501.13241