D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Krishna, Satyapriya, Zou, Andy, Gupta, Rahul, Jones, Eliot Krzysztof, Winter, Nick, Hendrycks, Dan, Kolter, J. Zico, Fredrikson, Matt, Matsoukas, Spyros
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912599621238784
author Krishna, Satyapriya
Zou, Andy
Gupta, Rahul
Jones, Eliot Krzysztof
Winter, Nick
Hendrycks, Dan
Kolter, J. Zico
Fredrikson, Matt
Matsoukas, Spyros
author_facet Krishna, Satyapriya
Zou, Andy
Gupta, Rahul
Jones, Eliot Krzysztof
Winter, Nick
Hendrycks, Dan
Kolter, J. Zico
Fredrikson, Matt
Matsoukas, Spyros
contents The safety and alignment of Large Language Models (LLMs) are critical for their responsible deployment. Current evaluation methods predominantly focus on identifying and preventing overtly harmful outputs. However, they often fail to address a more insidious failure mode: models that produce benign-appearing outputs while operating on malicious or deceptive internal reasoning. This vulnerability, often triggered by sophisticated system prompt injections, allows models to bypass conventional safety filters, posing a significant, underexplored risk. To address this gap, we introduce the Deceptive Reasoning Exposure Suite (D-REX), a novel dataset designed to evaluate the discrepancy between a model's internal reasoning process and its final output. D-REX was constructed through a competitive red-teaming exercise where participants crafted adversarial system prompts to induce such deceptive behaviors. Each sample in D-REX contains the adversarial system prompt, an end-user's test query, the model's seemingly innocuous response, and, crucially, the model's internal chain-of-thought, which reveals the underlying malicious intent. Our benchmark facilitates a new, essential evaluation task: the detection of deceptive alignment. We demonstrate that D-REX presents a significant challenge for existing models and safety mechanisms, highlighting the urgent need for new techniques that scrutinize the internal processes of LLMs, not just their final outputs.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17938
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
Krishna, Satyapriya
Zou, Andy
Gupta, Rahul
Jones, Eliot Krzysztof
Winter, Nick
Hendrycks, Dan
Kolter, J. Zico
Fredrikson, Matt
Matsoukas, Spyros
Computation and Language
The safety and alignment of Large Language Models (LLMs) are critical for their responsible deployment. Current evaluation methods predominantly focus on identifying and preventing overtly harmful outputs. However, they often fail to address a more insidious failure mode: models that produce benign-appearing outputs while operating on malicious or deceptive internal reasoning. This vulnerability, often triggered by sophisticated system prompt injections, allows models to bypass conventional safety filters, posing a significant, underexplored risk. To address this gap, we introduce the Deceptive Reasoning Exposure Suite (D-REX), a novel dataset designed to evaluate the discrepancy between a model's internal reasoning process and its final output. D-REX was constructed through a competitive red-teaming exercise where participants crafted adversarial system prompts to induce such deceptive behaviors. Each sample in D-REX contains the adversarial system prompt, an end-user's test query, the model's seemingly innocuous response, and, crucially, the model's internal chain-of-thought, which reveals the underlying malicious intent. Our benchmark facilitates a new, essential evaluation task: the detection of deceptive alignment. We demonstrate that D-REX presents a significant challenge for existing models and safety mechanisms, highlighting the urgent need for new techniques that scrutinize the internal processes of LLMs, not just their final outputs.
title D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2509.17938