Leveraging Domain Knowledge for Efficient Reward Modelling in RLHF: A Case-Study in E-Commerce Opinion Summarization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nath, Swaroop, Siledar, Tejpalsingh, Muddu, Sankara Sri Raghava Ravindra, Rangaraju, Rupasai, Khadilkar, Harshad, Bhattacharyya, Pushpak, Banerjee, Suman, Patil, Amey, Singh, Sudhanshu Shekhar, Chelliah, Muthusamy, Garera, Nikesh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916211285032960
author Nath, Swaroop
Siledar, Tejpalsingh
Muddu, Sankara Sri Raghava Ravindra
Rangaraju, Rupasai
Khadilkar, Harshad
Bhattacharyya, Pushpak
Banerjee, Suman
Patil, Amey
Singh, Sudhanshu Shekhar
Chelliah, Muthusamy
Garera, Nikesh
author_facet Nath, Swaroop
Siledar, Tejpalsingh
Muddu, Sankara Sri Raghava Ravindra
Rangaraju, Rupasai
Khadilkar, Harshad
Bhattacharyya, Pushpak
Banerjee, Suman
Patil, Amey
Singh, Sudhanshu Shekhar
Chelliah, Muthusamy
Garera, Nikesh
contents Reinforcement Learning from Human Feedback (RLHF) has become a dominating strategy in aligning Language Models (LMs) with human values/goals. The key to the strategy is learning a reward model ($φ$), which can reflect the latent reward model of humans. While this strategy has proven effective, the training methodology requires a lot of human preference annotation (usually in the order of tens of thousands) to train $φ$. Such a large-scale annotation is justifiable when it's a one-time effort, and the reward model is universally applicable. However, human goals are subjective and depend on the task, requiring task-specific preference annotations, which can be impractical to fulfill. To address this challenge, we propose a novel approach to infuse domain knowledge into $φ$, which reduces the amount of preference annotation required ($21\times$), omits Alignment Tax, and provides some interpretability. We validate our approach in E-Commerce Opinion Summarization, with a significant reduction in dataset size (to just $940$ samples) while advancing the SOTA ($\sim4$ point ROUGE-L improvement, $68\%$ of times preferred by humans over SOTA). Our contributions include a novel Reward Modeling technique and two new datasets: PromptOpinSumm (supervised data for Opinion Summarization) and OpinPref (a gold-standard human preference dataset). The proposed methodology opens up avenues for efficient RLHF, making it more adaptable to applications with varying human values. We release the artifacts (Code: github.com/efficient-rlhf. PromptOpinSumm: hf.co/prompt-opin-summ. OpinPref: hf.co/opin-pref) for usage under MIT License.
format Preprint
id arxiv_https___arxiv_org_abs_2402_15473
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Leveraging Domain Knowledge for Efficient Reward Modelling in RLHF: A Case-Study in E-Commerce Opinion Summarization
Nath, Swaroop
Siledar, Tejpalsingh
Muddu, Sankara Sri Raghava Ravindra
Rangaraju, Rupasai
Khadilkar, Harshad
Bhattacharyya, Pushpak
Banerjee, Suman
Patil, Amey
Singh, Sudhanshu Shekhar
Chelliah, Muthusamy
Garera, Nikesh
Computation and Language
Machine Learning
Reinforcement Learning from Human Feedback (RLHF) has become a dominating strategy in aligning Language Models (LMs) with human values/goals. The key to the strategy is learning a reward model ($φ$), which can reflect the latent reward model of humans. While this strategy has proven effective, the training methodology requires a lot of human preference annotation (usually in the order of tens of thousands) to train $φ$. Such a large-scale annotation is justifiable when it's a one-time effort, and the reward model is universally applicable. However, human goals are subjective and depend on the task, requiring task-specific preference annotations, which can be impractical to fulfill. To address this challenge, we propose a novel approach to infuse domain knowledge into $φ$, which reduces the amount of preference annotation required ($21\times$), omits Alignment Tax, and provides some interpretability. We validate our approach in E-Commerce Opinion Summarization, with a significant reduction in dataset size (to just $940$ samples) while advancing the SOTA ($\sim4$ point ROUGE-L improvement, $68\%$ of times preferred by humans over SOTA). Our contributions include a novel Reward Modeling technique and two new datasets: PromptOpinSumm (supervised data for Opinion Summarization) and OpinPref (a gold-standard human preference dataset). The proposed methodology opens up avenues for efficient RLHF, making it more adaptable to applications with varying human values. We release the artifacts (Code: github.com/efficient-rlhf. PromptOpinSumm: hf.co/prompt-opin-summ. OpinPref: hf.co/opin-pref) for usage under MIT License.
title Leveraging Domain Knowledge for Efficient Reward Modelling in RLHF: A Case-Study in E-Commerce Opinion Summarization
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2402.15473