Improved Off-policy Reinforcement Learning in Biological Sequence Design

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Hyeonah, Kim, Minsu, Yun, Taeyoung, Choi, Sanghyeok, Bengio, Emmanuel, Hernández-García, Alex, Park, Jinkyoo
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912433084301312
author Kim, Hyeonah
Kim, Minsu
Yun, Taeyoung
Choi, Sanghyeok
Bengio, Emmanuel
Hernández-García, Alex
Park, Jinkyoo
author_facet Kim, Hyeonah
Kim, Minsu
Yun, Taeyoung
Choi, Sanghyeok
Bengio, Emmanuel
Hernández-García, Alex
Park, Jinkyoo
contents Designing biological sequences with desired properties is challenging due to vast search spaces and limited evaluation budgets. Although reinforcement learning methods use proxy models for rapid reward evaluation, insufficient training data can cause proxy misspecification on out-of-distribution inputs. To address this, we propose a novel off-policy search, $δ$-Conservative Search, that enhances robustness by restricting policy exploration to reliable regions. Starting from high-score offline sequences, we inject noise by randomly masking tokens with probability $δ$, then denoise them using our policy. We further adapt $δ$ based on proxy uncertainty on each data point, aligning the level of conservativeness with model confidence. Experimental results show that our conservative search consistently enhances the off-policy training, outperforming existing machine learning methods in discovering high-score sequences across diverse tasks, including DNA, RNA, protein, and peptide design.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04461
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improved Off-policy Reinforcement Learning in Biological Sequence Design
Kim, Hyeonah
Kim, Minsu
Yun, Taeyoung
Choi, Sanghyeok
Bengio, Emmanuel
Hernández-García, Alex
Park, Jinkyoo
Machine Learning
Biomolecules
Designing biological sequences with desired properties is challenging due to vast search spaces and limited evaluation budgets. Although reinforcement learning methods use proxy models for rapid reward evaluation, insufficient training data can cause proxy misspecification on out-of-distribution inputs. To address this, we propose a novel off-policy search, $δ$-Conservative Search, that enhances robustness by restricting policy exploration to reliable regions. Starting from high-score offline sequences, we inject noise by randomly masking tokens with probability $δ$, then denoise them using our policy. We further adapt $δ$ based on proxy uncertainty on each data point, aligning the level of conservativeness with model confidence. Experimental results show that our conservative search consistently enhances the off-policy training, outperforming existing machine learning methods in discovering high-score sequences across diverse tasks, including DNA, RNA, protein, and peptide design.
title Improved Off-policy Reinforcement Learning in Biological Sequence Design
topic Machine Learning
Biomolecules
url https://arxiv.org/abs/2410.04461