Factored Diffusion Policies:Compositionally Generalized Robot Control with a Single Score Network

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mitra, Sayan, Yuceel, Ege, Giles, Noah, Pai, Abhishek
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916036901601280
author Mitra, Sayan
Yuceel, Ege
Giles, Noah
Pai, Abhishek
author_facet Mitra, Sayan
Yuceel, Ege
Giles, Noah
Pai, Abhishek
contents Robotic tasks are typically specified by a tuple of factors, such as the object to be grasped, the obstacles to be avoided, the color of the target, and so on. Collecting expert demonstrations for every combination of factor values grows combinatorially. We present factored diffusion policies: a single shared diffusion network trained with per-factor null-token dropout, whose score decomposes additively across factors at inference. Under approximate conditional independence between factors given the action-observation pair, this composition approximates the true joint score with a bounded uniform error, reducing the training-task budget from a product of factor cardinalities to a sum. A trajectory-tube certificate chains this score-level bound through the reverse-time sampling ODE and a contracting tracking controller into a closed-loop state-trajectory tube whose radius factors into an ODE-sensitivity constant and a per-factor score-error budget. Unlike compositional-diffusion methods for control that combine separately trained networks, we use one shared network. Drone racing experiments confirm both the generalization bound and the certificate. On state-based multi-gate racing, the factored policy passes 90% of held-out gates -- matching an oracle -- while a K-network composition baseline collapses to 3%; on vision-based single-gate traversal, it transfers zero-shot to an unseen venue with +11.7pp success-rate gain and 2.4X crash-rate reduction.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22596
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Factored Diffusion Policies:Compositionally Generalized Robot Control with a Single Score Network
Mitra, Sayan
Yuceel, Ege
Giles, Noah
Pai, Abhishek
Machine Learning
cs.LG
I.2.9; I.2.6; I.2.8
Robotic tasks are typically specified by a tuple of factors, such as the object to be grasped, the obstacles to be avoided, the color of the target, and so on. Collecting expert demonstrations for every combination of factor values grows combinatorially. We present factored diffusion policies: a single shared diffusion network trained with per-factor null-token dropout, whose score decomposes additively across factors at inference. Under approximate conditional independence between factors given the action-observation pair, this composition approximates the true joint score with a bounded uniform error, reducing the training-task budget from a product of factor cardinalities to a sum. A trajectory-tube certificate chains this score-level bound through the reverse-time sampling ODE and a contracting tracking controller into a closed-loop state-trajectory tube whose radius factors into an ODE-sensitivity constant and a per-factor score-error budget. Unlike compositional-diffusion methods for control that combine separately trained networks, we use one shared network. Drone racing experiments confirm both the generalization bound and the certificate. On state-based multi-gate racing, the factored policy passes 90% of held-out gates -- matching an oracle -- while a K-network composition baseline collapses to 3%; on vision-based single-gate traversal, it transfers zero-shot to an unseen venue with +11.7pp success-rate gain and 2.4X crash-rate reduction.
title Factored Diffusion Policies:Compositionally Generalized Robot Control with a Single Score Network
topic Machine Learning
cs.LG
I.2.9; I.2.6; I.2.8
url https://arxiv.org/abs/2605.22596