Red Teaming Language Models for Processing Contradictory Dialogues

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wen, Xiaofei, Li, Bangzheng, Huang, Tenghao, Chen, Muhao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914965076574208
author Wen, Xiaofei
Li, Bangzheng
Huang, Tenghao
Chen, Muhao
author_facet Wen, Xiaofei
Li, Bangzheng
Huang, Tenghao
Chen, Muhao
contents Most language models currently available are prone to self-contradiction during dialogues. To mitigate this issue, this study explores a novel contradictory dialogue processing task that aims to detect and modify contradictory statements in a conversation. This task is inspired by research on context faithfulness and dialogue comprehension, which have demonstrated that the detection and understanding of contradictions often necessitate detailed explanations. We develop a dataset comprising contradictory dialogues, in which one side of the conversation contradicts itself. Each dialogue is accompanied by an explanatory label that highlights the location and details of the contradiction. With this dataset, we present a Red Teaming framework for contradictory dialogue processing. The framework detects and attempts to explain the dialogue, then modifies the existing contradictory content using the explanation. Our experiments demonstrate that the framework improves the ability to detect contradictory dialogues and provides valid explanations. Additionally, it showcases distinct capabilities for modifying such dialogues. Our study highlights the importance of the logical inconsistency problem in conversational AI.
format Preprint
id arxiv_https___arxiv_org_abs_2405_10128
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Red Teaming Language Models for Processing Contradictory Dialogues
Wen, Xiaofei
Li, Bangzheng
Huang, Tenghao
Chen, Muhao
Computation and Language
Artificial Intelligence
Most language models currently available are prone to self-contradiction during dialogues. To mitigate this issue, this study explores a novel contradictory dialogue processing task that aims to detect and modify contradictory statements in a conversation. This task is inspired by research on context faithfulness and dialogue comprehension, which have demonstrated that the detection and understanding of contradictions often necessitate detailed explanations. We develop a dataset comprising contradictory dialogues, in which one side of the conversation contradicts itself. Each dialogue is accompanied by an explanatory label that highlights the location and details of the contradiction. With this dataset, we present a Red Teaming framework for contradictory dialogue processing. The framework detects and attempts to explain the dialogue, then modifies the existing contradictory content using the explanation. Our experiments demonstrate that the framework improves the ability to detect contradictory dialogues and provides valid explanations. Additionally, it showcases distinct capabilities for modifying such dialogues. Our study highlights the importance of the logical inconsistency problem in conversational AI.
title Red Teaming Language Models for Processing Contradictory Dialogues
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2405.10128