Steering Masked Discrete Diffusion Models via Discrete Denoising Posterior Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rector-Brooks, Jarrid, Hasan, Mohsin, Peng, Zhangzhi, Quinn, Zachary, Liu, Chenghao, Mittal, Sarthak, Dziri, Nouha, Bronstein, Michael, Bengio, Yoshua, Chatterjee, Pranam, Tong, Alexander, Bose, Avishek Joey
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929535785631744
author Rector-Brooks, Jarrid
Hasan, Mohsin
Peng, Zhangzhi
Quinn, Zachary
Liu, Chenghao
Mittal, Sarthak
Dziri, Nouha
Bronstein, Michael
Bengio, Yoshua
Chatterjee, Pranam
Tong, Alexander
Bose, Avishek Joey
author_facet Rector-Brooks, Jarrid
Hasan, Mohsin
Peng, Zhangzhi
Quinn, Zachary
Liu, Chenghao
Mittal, Sarthak
Dziri, Nouha
Bronstein, Michael
Bengio, Yoshua
Chatterjee, Pranam
Tong, Alexander
Bose, Avishek Joey
contents Generative modeling of discrete data underlies important applications spanning text-based agents like ChatGPT to the design of the very building blocks of life in protein sequences. However, application domains need to exert control over the generated data by steering the generative process - typically via RLHF - to satisfy a specified property, reward, or affinity metric. In this paper, we study the problem of steering Masked Diffusion Models (MDMs), a recent class of discrete diffusion models that offer a compelling alternative to traditional autoregressive models. We introduce Discrete Denoising Posterior Prediction (DDPP), a novel framework that casts the task of steering pre-trained MDMs as a problem of probabilistic inference by learning to sample from a target Bayesian posterior. Our DDPP framework leads to a family of three novel objectives that are all simulation-free, and thus scalable while applying to general non-differentiable reward functions. Empirically, we instantiate DDPP by steering MDMs to perform class-conditional pixel-level image modeling, RLHF-based alignment of MDMs using text-based rewards, and finetuning protein language models to generate more diverse secondary structures and shorter proteins. We substantiate our designs via wet-lab validation, where we observe transient expression of reward-optimized protein sequences.
format Preprint
id arxiv_https___arxiv_org_abs_2410_08134
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Steering Masked Discrete Diffusion Models via Discrete Denoising Posterior Prediction
Rector-Brooks, Jarrid
Hasan, Mohsin
Peng, Zhangzhi
Quinn, Zachary
Liu, Chenghao
Mittal, Sarthak
Dziri, Nouha
Bronstein, Michael
Bengio, Yoshua
Chatterjee, Pranam
Tong, Alexander
Bose, Avishek Joey
Machine Learning
Artificial Intelligence
Generative modeling of discrete data underlies important applications spanning text-based agents like ChatGPT to the design of the very building blocks of life in protein sequences. However, application domains need to exert control over the generated data by steering the generative process - typically via RLHF - to satisfy a specified property, reward, or affinity metric. In this paper, we study the problem of steering Masked Diffusion Models (MDMs), a recent class of discrete diffusion models that offer a compelling alternative to traditional autoregressive models. We introduce Discrete Denoising Posterior Prediction (DDPP), a novel framework that casts the task of steering pre-trained MDMs as a problem of probabilistic inference by learning to sample from a target Bayesian posterior. Our DDPP framework leads to a family of three novel objectives that are all simulation-free, and thus scalable while applying to general non-differentiable reward functions. Empirically, we instantiate DDPP by steering MDMs to perform class-conditional pixel-level image modeling, RLHF-based alignment of MDMs using text-based rewards, and finetuning protein language models to generate more diverse secondary structures and shorter proteins. We substantiate our designs via wet-lab validation, where we observe transient expression of reward-optimized protein sequences.
title Steering Masked Discrete Diffusion Models via Discrete Denoising Posterior Prediction
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.08134