Discrete Tilt Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yuyuan, Wang, Shiyi, Potaptchik, Peter, Kim, Jaeyeon, Albergo, Michael S.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910233848184832
author Chen, Yuyuan
Wang, Shiyi
Potaptchik, Peter
Kim, Jaeyeon
Albergo, Michael S.
author_facet Chen, Yuyuan
Wang, Shiyi
Potaptchik, Peter
Kim, Jaeyeon
Albergo, Michael S.
contents Masked diffusion large language models (dLLMs) are a promising alternative to autoregressive generation. While reinforcement learning (RL) methods have recently been adapted to dLLM fine-tuning, their objectives typically depend on sequence-level marginal likelihoods, which are intractable for masked diffusion models. To address this, we derive Discrete Tilt Matching (DTM), a likelihood-free method that recasts dLLM fine-tuning as state-level matching of local unmasking posteriors under reward tilting. DTM takes the form of a weighted cross-entropy objective with explicit minimizer, and admits control variates that improve training stability. On a synthetic maze-planning task, we analyze how DTM's annealing schedule and control variates affect training stability and prevent mode collapse. At scale, fine-tuning LLaDA-8B-Instruct with DTM yields strong gains on Sudoku and Countdown while remaining competitive on MATH500 and GSM8K.
format Preprint
id arxiv_https___arxiv_org_abs_2604_18739
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Discrete Tilt Matching
Chen, Yuyuan
Wang, Shiyi
Potaptchik, Peter
Kim, Jaeyeon
Albergo, Michael S.
Machine Learning
Masked diffusion large language models (dLLMs) are a promising alternative to autoregressive generation. While reinforcement learning (RL) methods have recently been adapted to dLLM fine-tuning, their objectives typically depend on sequence-level marginal likelihoods, which are intractable for masked diffusion models. To address this, we derive Discrete Tilt Matching (DTM), a likelihood-free method that recasts dLLM fine-tuning as state-level matching of local unmasking posteriors under reward tilting. DTM takes the form of a weighted cross-entropy objective with explicit minimizer, and admits control variates that improve training stability. On a synthetic maze-planning task, we analyze how DTM's annealing schedule and control variates affect training stability and prevent mode collapse. At scale, fine-tuning LLaDA-8B-Instruct with DTM yields strong gains on Sudoku and Countdown while remaining competitive on MATH500 and GSM8K.
title Discrete Tilt Matching
topic Machine Learning
url https://arxiv.org/abs/2604.18739