Saved in:
Bibliographic Details
Main Authors: Wang, Xiaohua, Huang, Zisu, Zhang, Feiran, Xu, Zhibo, Zhang, Cenyuan, Qian, Qi, Zheng, Xiaoqing, Huang, Xuanjing
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2407.01461
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908425838919680
author Wang, Xiaohua
Huang, Zisu
Zhang, Feiran
Xu, Zhibo
Zhang, Cenyuan
Qian, Qi
Zheng, Xiaoqing
Huang, Xuanjing
author_facet Wang, Xiaohua
Huang, Zisu
Zhang, Feiran
Xu, Zhibo
Zhang, Cenyuan
Qian, Qi
Zheng, Xiaoqing
Huang, Xuanjing
contents The capacity of large language models (LLMs) to generate honest, harmless, and helpful responses heavily relies on the quality of user prompts. However, these prompts often tend to be brief and vague, thereby significantly limiting the full potential of LLMs. Moreover, harmful prompts can be meticulously crafted and manipulated by adversaries to jailbreak LLMs, inducing them to produce potentially toxic content. To enhance the capabilities of LLMs while maintaining strong robustness against harmful jailbreak inputs, this study proposes a transferable and pluggable framework that refines user prompts before they are input into LLMs. This strategy improves the quality of the queries, empowering LLMs to generate more truthful, benign and useful responses. Specifically, a lightweight query refinement model is introduced and trained using a specially designed reinforcement learning approach that incorporates multiple objectives to enhance particular capabilities of LLMs. Extensive experiments demonstrate that the refinement model not only improves the quality of responses but also strengthens their robustness against jailbreak attacks. Code is available at: https://github.com/Huangzisu/query-refinement .
format Preprint
id arxiv_https___arxiv_org_abs_2407_01461
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing the Capability and Robustness of Large Language Models through Reinforcement Learning-Driven Query Refinement
Wang, Xiaohua
Huang, Zisu
Zhang, Feiran
Xu, Zhibo
Zhang, Cenyuan
Qian, Qi
Zheng, Xiaoqing
Huang, Xuanjing
Computation and Language
The capacity of large language models (LLMs) to generate honest, harmless, and helpful responses heavily relies on the quality of user prompts. However, these prompts often tend to be brief and vague, thereby significantly limiting the full potential of LLMs. Moreover, harmful prompts can be meticulously crafted and manipulated by adversaries to jailbreak LLMs, inducing them to produce potentially toxic content. To enhance the capabilities of LLMs while maintaining strong robustness against harmful jailbreak inputs, this study proposes a transferable and pluggable framework that refines user prompts before they are input into LLMs. This strategy improves the quality of the queries, empowering LLMs to generate more truthful, benign and useful responses. Specifically, a lightweight query refinement model is introduced and trained using a specially designed reinforcement learning approach that incorporates multiple objectives to enhance particular capabilities of LLMs. Extensive experiments demonstrate that the refinement model not only improves the quality of responses but also strengthens their robustness against jailbreak attacks. Code is available at: https://github.com/Huangzisu/query-refinement .
title Enhancing the Capability and Robustness of Large Language Models through Reinforcement Learning-Driven Query Refinement
topic Computation and Language
url https://arxiv.org/abs/2407.01461