Tell Me What You Don't Know: Enhancing Refusal Capabilities of Role-Playing Agents via Representation Space Analysis and Editing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Wenhao, An, Siyu, Lu, Junru, Wu, Muling, Li, Tianlong, Wang, Xiaohua, lv, Changze, Zheng, Xiaoqing, Yin, Di, Sun, Xing, Huang, Xuanjing
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916792281071616
author Liu, Wenhao
An, Siyu
Lu, Junru
Wu, Muling
Li, Tianlong
Wang, Xiaohua
lv, Changze
Zheng, Xiaoqing
Yin, Di
Sun, Xing
Huang, Xuanjing
author_facet Liu, Wenhao
An, Siyu
Lu, Junru
Wu, Muling
Li, Tianlong
Wang, Xiaohua
lv, Changze
Zheng, Xiaoqing
Yin, Di
Sun, Xing
Huang, Xuanjing
contents Role-Playing Agents (RPAs) have shown remarkable performance in various applications, yet they often struggle to recognize and appropriately respond to hard queries that conflict with their role-play knowledge. To investigate RPAs' performance when faced with different types of conflicting requests, we develop an evaluation benchmark that includes contextual knowledge conflicting requests, parametric knowledge conflicting requests, and non-conflicting requests to assess RPAs' ability to identify conflicts and refuse to answer appropriately without over-refusing. Through extensive evaluation, we find that most RPAs behave significant performance gaps toward different conflict requests. To elucidate the reasons, we conduct an in-depth representation-level analysis of RPAs under various conflict scenarios. Our findings reveal the existence of rejection regions and direct response regions within the model's forwarding representation, and thus influence the RPA's final response behavior. Therefore, we introduce a lightweight representation editing approach that conveniently shifts conflicting requests to the rejection region, thereby enhancing the model's refusal accuracy. The experimental results validate the effectiveness of our editing method, improving RPAs' refusal ability of conflicting requests while maintaining their general role-playing capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2409_16913
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Tell Me What You Don't Know: Enhancing Refusal Capabilities of Role-Playing Agents via Representation Space Analysis and Editing
Liu, Wenhao
An, Siyu
Lu, Junru
Wu, Muling
Li, Tianlong
Wang, Xiaohua
lv, Changze
Zheng, Xiaoqing
Yin, Di
Sun, Xing
Huang, Xuanjing
Artificial Intelligence
Role-Playing Agents (RPAs) have shown remarkable performance in various applications, yet they often struggle to recognize and appropriately respond to hard queries that conflict with their role-play knowledge. To investigate RPAs' performance when faced with different types of conflicting requests, we develop an evaluation benchmark that includes contextual knowledge conflicting requests, parametric knowledge conflicting requests, and non-conflicting requests to assess RPAs' ability to identify conflicts and refuse to answer appropriately without over-refusing. Through extensive evaluation, we find that most RPAs behave significant performance gaps toward different conflict requests. To elucidate the reasons, we conduct an in-depth representation-level analysis of RPAs under various conflict scenarios. Our findings reveal the existence of rejection regions and direct response regions within the model's forwarding representation, and thus influence the RPA's final response behavior. Therefore, we introduce a lightweight representation editing approach that conveniently shifts conflicting requests to the rejection region, thereby enhancing the model's refusal accuracy. The experimental results validate the effectiveness of our editing method, improving RPAs' refusal ability of conflicting requests while maintaining their general role-playing capabilities.
title Tell Me What You Don't Know: Enhancing Refusal Capabilities of Role-Playing Agents via Representation Space Analysis and Editing
topic Artificial Intelligence
url https://arxiv.org/abs/2409.16913