Saved in:
Bibliographic Details
Main Authors: Chaudhary, Sapana, Dinesha, Ujwal, Kalathil, Dileep, Shakkottai, Srinivas
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2501.06911
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917891118465024
author Chaudhary, Sapana
Dinesha, Ujwal
Kalathil, Dileep
Shakkottai, Srinivas
author_facet Chaudhary, Sapana
Dinesha, Ujwal
Kalathil, Dileep
Shakkottai, Srinivas
contents We consider the challenge of mitigating the generation of negative or toxic content by the Large Language Models (LLMs) in response to certain prompts. We propose integrating risk-averse principles into LLM fine-tuning to minimize the occurrence of harmful outputs, particularly rare but significant events. By optimizing the risk measure of Conditional Value at Risk (CVaR), our methodology trains LLMs to exhibit superior performance in avoiding toxic outputs while maintaining effectiveness in generative tasks. Empirical evaluations on sentiment modification and toxicity mitigation tasks demonstrate the efficacy of risk-averse reinforcement learning with human feedback (RLHF) in promoting a safer and more constructive online discourse environment.
format Preprint
id arxiv_https___arxiv_org_abs_2501_06911
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Risk-Averse Finetuning of Large Language Models
Chaudhary, Sapana
Dinesha, Ujwal
Kalathil, Dileep
Shakkottai, Srinivas
Artificial Intelligence
Computation and Language
We consider the challenge of mitigating the generation of negative or toxic content by the Large Language Models (LLMs) in response to certain prompts. We propose integrating risk-averse principles into LLM fine-tuning to minimize the occurrence of harmful outputs, particularly rare but significant events. By optimizing the risk measure of Conditional Value at Risk (CVaR), our methodology trains LLMs to exhibit superior performance in avoiding toxic outputs while maintaining effectiveness in generative tasks. Empirical evaluations on sentiment modification and toxicity mitigation tasks demonstrate the efficacy of risk-averse reinforcement learning with human feedback (RLHF) in promoting a safer and more constructive online discourse environment.
title Risk-Averse Finetuning of Large Language Models
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2501.06911