Bootstrapped Training of Score-Conditioned Generator for Offline Design of Biological Sequences

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kim, Minsu, Berto, Federico, Ahn, Sungsoo, Park, Jinkyoo
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916172654444544
author Kim, Minsu
Berto, Federico
Ahn, Sungsoo
Park, Jinkyoo
author_facet Kim, Minsu
Berto, Federico
Ahn, Sungsoo
Park, Jinkyoo
contents We study the problem of optimizing biological sequences, e.g., proteins, DNA, and RNA, to maximize a black-box score function that is only evaluated in an offline dataset. We propose a novel solution, bootstrapped training of score-conditioned generator (BootGen) algorithm. Our algorithm repeats a two-stage process. In the first stage, our algorithm trains the biological sequence generator with rank-based weights to enhance the accuracy of sequence generation based on high scores. The subsequent stage involves bootstrapping, which augments the training dataset with self-generated data labeled by a proxy score function. Our key idea is to align the score-based generation with a proxy score function, which distills the knowledge of the proxy score function to the generator. After training, we aggregate samples from multiple bootstrapped generators and proxies to produce a diverse design. Extensive experiments show that our method outperforms competitive baselines on biological sequential design tasks. We provide reproducible source code: \href{https://github.com/kaist-silab/bootgen}{https://github.com/kaist-silab/bootgen}.
format Preprint
id arxiv_https___arxiv_org_abs_2306_03111
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Bootstrapped Training of Score-Conditioned Generator for Offline Design of Biological Sequences
Kim, Minsu
Berto, Federico
Ahn, Sungsoo
Park, Jinkyoo
Quantitative Methods
Machine Learning
We study the problem of optimizing biological sequences, e.g., proteins, DNA, and RNA, to maximize a black-box score function that is only evaluated in an offline dataset. We propose a novel solution, bootstrapped training of score-conditioned generator (BootGen) algorithm. Our algorithm repeats a two-stage process. In the first stage, our algorithm trains the biological sequence generator with rank-based weights to enhance the accuracy of sequence generation based on high scores. The subsequent stage involves bootstrapping, which augments the training dataset with self-generated data labeled by a proxy score function. Our key idea is to align the score-based generation with a proxy score function, which distills the knowledge of the proxy score function to the generator. After training, we aggregate samples from multiple bootstrapped generators and proxies to produce a diverse design. Extensive experiments show that our method outperforms competitive baselines on biological sequential design tasks. We provide reproducible source code: \href{https://github.com/kaist-silab/bootgen}{https://github.com/kaist-silab/bootgen}.
title Bootstrapped Training of Score-Conditioned Generator for Offline Design of Biological Sequences
topic Quantitative Methods
Machine Learning
url https://arxiv.org/abs/2306.03111