Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection Generation

Min, Do June; Perez-Rosas, Veronica; Resnicow, Ken; Mihalcea, Rada

Citation Details

In this paper, we study the problem of multi-reward reinforcement learning to jointly optimize for multiple text qualities for natural language generation. We focus on the task of counselor reflection generation, where we optimize the generators to simultaneously improve the fluency, coherence, and reflection quality of generated counselor responses. We introduce two novel bandit methods, DynaOpt and C-DynaOpt, which rely on the broad strategy of combining rewards into a single value and optimizing them simultaneously. Specifically, we employ non-contextual and contextual multi-arm bandits to dynamically adjust multiple reward weights during training. Through automatic and manual evaluations, we show that our proposed techniques, DynaOpt and C-DynaOpt, outperform existing naive and bandit baselines, showcasing their potential for enhancing language models. more »

Award ID(s):: 2306372

PAR ID:: 10616294

Author(s) / Creator(s):: Min, Do June; Perez-Rosas, Veronica; Resnicow, Ken; Mihalcea, Rada

Publisher / Repository:: ELRA and ICCL

Date Published:: 2024-08-01

Format(s):: Medium: X

Location:: Torino, Italia

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
The DOI is not currently available.

More Like this