Training-Free Self-Correction for Multimodal Masked Diffusion Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ouyang, Yidong, Hu, Panwen, Wan, Zhengyan, Wang, Zhe, Xie, Liyan, Bespalov, Dmitriy, Wu, Ying Nian, Cheng, Guang, Zha, Hongyuan, Sun, Qiang
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910009863962624
author Ouyang, Yidong
Hu, Panwen
Wan, Zhengyan
Wang, Zhe
Xie, Liyan
Bespalov, Dmitriy
Wu, Ying Nian
Cheng, Guang
Zha, Hongyuan
Sun, Qiang
author_facet Ouyang, Yidong
Hu, Panwen
Wan, Zhengyan
Wang, Zhe
Xie, Liyan
Bespalov, Dmitriy
Wu, Ying Nian
Cheng, Guang
Zha, Hongyuan
Sun, Qiang
contents Masked diffusion models have emerged as a powerful framework for text and multimodal generation. However, their sampling procedure updates multiple tokens simultaneously and treats generated tokens as immutable, which may lead to error accumulation when early mistakes cannot be revised. In this work, we revisit existing self-correction methods and identify limitations stemming from additional training requirements or reliance on misaligned likelihood estimates. We propose a training-free self-correction framework that exploits the inductive biases of pre-trained masked diffusion models. Without modifying model parameters or introducing auxiliary evaluators, our method significantly improves generation quality on text-to-image generation and multimodal understanding tasks with reduced sampling steps. Moreover, the proposed framework generalizes across different masked diffusion architectures, highlighting its robustness and practical applicability. Code can be found in https://github.com/huge123/FreeCorrection.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02927
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Training-Free Self-Correction for Multimodal Masked Diffusion Models
Ouyang, Yidong
Hu, Panwen
Wan, Zhengyan
Wang, Zhe
Xie, Liyan
Bespalov, Dmitriy
Wu, Ying Nian
Cheng, Guang
Zha, Hongyuan
Sun, Qiang
Machine Learning
Masked diffusion models have emerged as a powerful framework for text and multimodal generation. However, their sampling procedure updates multiple tokens simultaneously and treats generated tokens as immutable, which may lead to error accumulation when early mistakes cannot be revised. In this work, we revisit existing self-correction methods and identify limitations stemming from additional training requirements or reliance on misaligned likelihood estimates. We propose a training-free self-correction framework that exploits the inductive biases of pre-trained masked diffusion models. Without modifying model parameters or introducing auxiliary evaluators, our method significantly improves generation quality on text-to-image generation and multimodal understanding tasks with reduced sampling steps. Moreover, the proposed framework generalizes across different masked diffusion architectures, highlighting its robustness and practical applicability. Code can be found in https://github.com/huge123/FreeCorrection.
title Training-Free Self-Correction for Multimodal Masked Diffusion Models
topic Machine Learning
url https://arxiv.org/abs/2602.02927