Discrete Compositional Generation via General Soft Operators and Robust Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiralerspong, Marco, Derman, Esther, Vucetic, Danilo, Malkin, Nikolay, Sun, Bilun, Zhang, Tianyu, Bacon, Pierre-Luc, Gidel, Gauthier
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908585414361088
author Jiralerspong, Marco
Derman, Esther
Vucetic, Danilo
Malkin, Nikolay
Sun, Bilun
Zhang, Tianyu
Bacon, Pierre-Luc
Gidel, Gauthier
author_facet Jiralerspong, Marco
Derman, Esther
Vucetic, Danilo
Malkin, Nikolay
Sun, Bilun
Zhang, Tianyu
Bacon, Pierre-Luc
Gidel, Gauthier
contents A major bottleneck in scientific discovery consists of narrowing an exponentially large set of objects, such as proteins or molecules, to a small set of promising candidates with desirable properties. While this process can rely on expert knowledge, recent methods leverage reinforcement learning (RL) guided by a proxy reward function to enable this filtering. By employing various forms of entropy regularization, these methods aim to learn samplers that generate diverse candidates that are highly rated by the proxy function. In this work, we make two main contributions. First, we show that these methods are liable to generate overly diverse, suboptimal candidates in large search spaces. To address this issue, we introduce a novel unified operator that combines several regularized RL operators into a general framework that better targets peakier sampling distributions. Secondly, we offer a novel, robust RL perspective of this filtering process. The regularization can be interpreted as robustness to a compositional form of uncertainty in the proxy function (i.e., the true evaluation of a candidate differs from the proxy's evaluation). Our analysis leads us to a novel, easy-to-use algorithm we name trajectory general mellowmax (TGM): we show it identifies higher quality, diverse candidates than baselines in both synthetic and real-world tasks. Code: https://github.com/marcojira/tgm.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17007
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Discrete Compositional Generation via General Soft Operators and Robust Reinforcement Learning
Jiralerspong, Marco
Derman, Esther
Vucetic, Danilo
Malkin, Nikolay
Sun, Bilun
Zhang, Tianyu
Bacon, Pierre-Luc
Gidel, Gauthier
Machine Learning
A major bottleneck in scientific discovery consists of narrowing an exponentially large set of objects, such as proteins or molecules, to a small set of promising candidates with desirable properties. While this process can rely on expert knowledge, recent methods leverage reinforcement learning (RL) guided by a proxy reward function to enable this filtering. By employing various forms of entropy regularization, these methods aim to learn samplers that generate diverse candidates that are highly rated by the proxy function. In this work, we make two main contributions. First, we show that these methods are liable to generate overly diverse, suboptimal candidates in large search spaces. To address this issue, we introduce a novel unified operator that combines several regularized RL operators into a general framework that better targets peakier sampling distributions. Secondly, we offer a novel, robust RL perspective of this filtering process. The regularization can be interpreted as robustness to a compositional form of uncertainty in the proxy function (i.e., the true evaluation of a candidate differs from the proxy's evaluation). Our analysis leads us to a novel, easy-to-use algorithm we name trajectory general mellowmax (TGM): we show it identifies higher quality, diverse candidates than baselines in both synthetic and real-world tasks. Code: https://github.com/marcojira/tgm.
title Discrete Compositional Generation via General Soft Operators and Robust Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2506.17007