DiscreteRTC: Discrete Diffusion Policies are Natural Asynchronous Executors

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Pengcheng, Hong, Kaiwen, Peng, Chensheng, Driggs-Campbell, Katherine, Tomizuka, Masayoshi, Xu, Chenfeng, Tang, Chen
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918470771277824
author Wang, Pengcheng
Hong, Kaiwen
Peng, Chensheng
Driggs-Campbell, Katherine
Tomizuka, Masayoshi
Xu, Chenfeng
Tang, Chen
author_facet Wang, Pengcheng
Hong, Kaiwen
Peng, Chensheng
Driggs-Campbell, Katherine
Tomizuka, Masayoshi
Xu, Chenfeng
Tang, Chen
contents Unlike chatbots, physical AI must act while the world keeps evolving. Therefore, the inter-chunk pause of synchronous executors are fatal for dynamic tasks regardless of how fast the inference is. Asynchronous execution -- thinking while acting -- is therefore a structural requirement, and real-time chunking (RTC) makes it viable by recasting chunk transitions as inpainting: freezing committed actions and consistently generating the remainder. However, RTC with flow-matching policy is structurally suboptimal: its inpainting comes from inference-time corrections rather than the base policy, yielding little pre-training benefit, specific fine-tuning, heuristic guidance, and extra computation that inflates the latency. In this work, we observe that discrete diffusion policies, which generate actions by iteratively unmasking, are natural asynchronous executors that resolve all limitations at once: they are fine-tuning free since inpainting is their native operation, while early stopping further provides adaptive guidance and reduces inference cost. We propose DiscreteRTC, which replaces external corrections with native unmasking, and show on dynamic simulated benchmarks and real-world dynamic manipulation tasks that it achieves higher success rates than continuous RTC and other baselines. In summary, DiscreteRTC is simpler to implement with 0 lines of code for async inpainting, faster at inference with only 0.7x computation compared with generating actions from scratch, and better at execution with 50% higher success rate in real-world dynamic pick task compared with flow-matching-based RTC. More visualizations are on https://outsider86.github.io/DiscreteRTCSite/.
format Preprint
id arxiv_https___arxiv_org_abs_2604_25050
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DiscreteRTC: Discrete Diffusion Policies are Natural Asynchronous Executors
Wang, Pengcheng
Hong, Kaiwen
Peng, Chensheng
Driggs-Campbell, Katherine
Tomizuka, Masayoshi
Xu, Chenfeng
Tang, Chen
Robotics
Unlike chatbots, physical AI must act while the world keeps evolving. Therefore, the inter-chunk pause of synchronous executors are fatal for dynamic tasks regardless of how fast the inference is. Asynchronous execution -- thinking while acting -- is therefore a structural requirement, and real-time chunking (RTC) makes it viable by recasting chunk transitions as inpainting: freezing committed actions and consistently generating the remainder. However, RTC with flow-matching policy is structurally suboptimal: its inpainting comes from inference-time corrections rather than the base policy, yielding little pre-training benefit, specific fine-tuning, heuristic guidance, and extra computation that inflates the latency. In this work, we observe that discrete diffusion policies, which generate actions by iteratively unmasking, are natural asynchronous executors that resolve all limitations at once: they are fine-tuning free since inpainting is their native operation, while early stopping further provides adaptive guidance and reduces inference cost. We propose DiscreteRTC, which replaces external corrections with native unmasking, and show on dynamic simulated benchmarks and real-world dynamic manipulation tasks that it achieves higher success rates than continuous RTC and other baselines. In summary, DiscreteRTC is simpler to implement with 0 lines of code for async inpainting, faster at inference with only 0.7x computation compared with generating actions from scratch, and better at execution with 50% higher success rate in real-world dynamic pick task compared with flow-matching-based RTC. More visualizations are on https://outsider86.github.io/DiscreteRTCSite/.
title DiscreteRTC: Discrete Diffusion Policies are Natural Asynchronous Executors
topic Robotics
url https://arxiv.org/abs/2604.25050