Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Min, Do June, Perez-Rosas, Veronica, Resnicow, Kenneth, Mihalcea, Rada
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916168237842432
author Min, Do June
Perez-Rosas, Veronica
Resnicow, Kenneth
Mihalcea, Rada
author_facet Min, Do June
Perez-Rosas, Veronica
Resnicow, Kenneth
Mihalcea, Rada
contents In this paper, we study the problem of multi-reward reinforcement learning to jointly optimize for multiple text qualities for natural language generation. We focus on the task of counselor reflection generation, where we optimize the generators to simultaneously improve the fluency, coherence, and reflection quality of generated counselor responses. We introduce two novel bandit methods, DynaOpt and C-DynaOpt, which rely on the broad strategy of combining rewards into a single value and optimizing them simultaneously. Specifically, we employ non-contextual and contextual multi-arm bandits to dynamically adjust multiple reward weights during training. Through automatic and manual evaluations, we show that our proposed techniques, DynaOpt and C-DynaOpt, outperform existing naive and bandit baselines, showcasing their potential for enhancing language models.
format Preprint
id arxiv_https___arxiv_org_abs_2403_13578
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection Generation
Min, Do June
Perez-Rosas, Veronica
Resnicow, Kenneth
Mihalcea, Rada
Computation and Language
Machine Learning
In this paper, we study the problem of multi-reward reinforcement learning to jointly optimize for multiple text qualities for natural language generation. We focus on the task of counselor reflection generation, where we optimize the generators to simultaneously improve the fluency, coherence, and reflection quality of generated counselor responses. We introduce two novel bandit methods, DynaOpt and C-DynaOpt, which rely on the broad strategy of combining rewards into a single value and optimizing them simultaneously. Specifically, we employ non-contextual and contextual multi-arm bandits to dynamically adjust multiple reward weights during training. Through automatic and manual evaluations, we show that our proposed techniques, DynaOpt and C-DynaOpt, outperform existing naive and bandit baselines, showcasing their potential for enhancing language models.
title Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection Generation
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2403.13578