D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Krishna, Satyapriya, Zou, Andy, Gupta, Rahul, Jones, Eliot Krzysztof, Winter, Nick, Hendrycks, Dan, Kolter, J. Zico, Fredrikson, Matt, Matsoukas, Spyros |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Language Model Reasoning about Confidential Information
by: Sam, Dylan, et al.
Published: (2025)
by: Sam, Dylan, et al.
Published: (2025)
Adversarial Attacks on Robotic Vision Language Action Models
by: Jones, Eliot Krzysztof, et al.
Published: (2025)
by: Jones, Eliot Krzysztof, et al.
Published: (2025)
Improving Alignment and Robustness with Circuit Breakers
by: Zou, Andy, et al.
Published: (2024)
by: Zou, Andy, et al.
Published: (2024)
Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
by: Zou, Andy, et al.
Published: (2025)
by: Zou, Andy, et al.
Published: (2025)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Evaluating Nova 2.0 Lite model under Amazon's Frontier Model Safety Framework
by: Krishna, Satyapriya, et al.
Published: (2026)
by: Krishna, Satyapriya, et al.
Published: (2026)
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
by: Lin, Justin W., et al.
Published: (2025)
by: Lin, Justin W., et al.
Published: (2025)
Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning
by: Huang, Benhao, et al.
Published: (2026)
by: Huang, Benhao, et al.
Published: (2026)
Safety Pretraining: Toward the Next Generation of Safe AI
by: Maini, Pratyush, et al.
Published: (2025)
by: Maini, Pratyush, et al.
Published: (2025)
TextQuests: How Good are LLMs at Text-Based Video Games?
by: Phan, Long, et al.
Published: (2025)
by: Phan, Long, et al.
Published: (2025)
Mimetic Initialization of MLPs
by: Trockman, Asher, et al.
Published: (2026)
by: Trockman, Asher, et al.
Published: (2026)
AcceleratedLiNGAM: Learning Causal DAGs at the speed of GPUs
by: Akinwande, Victor, et al.
Published: (2024)
by: Akinwande, Victor, et al.
Published: (2024)
Measuring Five-Nines Reliability: Sample-Efficient LLM Evaluation in Saturated Benchmarks
by: Kim, Eungyeup, et al.
Published: (2026)
by: Kim, Eungyeup, et al.
Published: (2026)
Introduction to AI Safety, Ethics, and Society
by: Hendrycks, Dan
Published: (2024)
by: Hendrycks, Dan
Published: (2024)
Introduction to AI Safety, Ethics, and Society
by: Hendrycks, Dan
Published: (2024)
by: Hendrycks, Dan
Published: (2024)
FUSE-ing Language Models: Zero-Shot Adapter Discovery for Prompt Optimization Across Tokenizers
by: Williams, Joshua Nathaniel, et al.
Published: (2024)
by: Williams, Joshua Nathaniel, et al.
Published: (2024)
Why is SAM Robust to Label Noise?
by: Baek, Christina, et al.
Published: (2024)
by: Baek, Christina, et al.
Published: (2024)
Finetuning CLIP to Reason about Pairwise Differences
by: Sam, Dylan, et al.
Published: (2024)
by: Sam, Dylan, et al.
Published: (2024)
Weight Ensembling Improves Reasoning in Language Models
by: Dang, Xingyu, et al.
Published: (2025)
by: Dang, Xingyu, et al.
Published: (2025)
Massive Activations in Large Language Models
by: Sun, Mingjie, et al.
Published: (2024)
by: Sun, Mingjie, et al.
Published: (2024)
A Simple and Effective Pruning Approach for Large Language Models
by: Sun, Mingjie, et al.
Published: (2023)
by: Sun, Mingjie, et al.
Published: (2023)
Predicting the Performance of Black-box LLMs through Follow-up Queries
by: Sam, Dylan, et al.
Published: (2025)
by: Sam, Dylan, et al.
Published: (2025)
One-Step Diffusion Distillation via Deep Equilibrium Models
by: Geng, Zhengyang, et al.
Published: (2023)
by: Geng, Zhengyang, et al.
Published: (2023)
Diffusing Differentiable Representations
by: Savani, Yash, et al.
Published: (2024)
by: Savani, Yash, et al.
Published: (2024)
Representation Engineering: A Top-Down Approach to AI Transparency
by: Zou, Andy, et al.
Published: (2023)
by: Zou, Andy, et al.
Published: (2023)
Compressed Sensing for Capability Localization in Large Language Models
by: Bair, Anna, et al.
Published: (2026)
by: Bair, Anna, et al.
Published: (2026)
V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions
by: Fan, Chenrui, et al.
Published: (2025)
by: Fan, Chenrui, et al.
Published: (2025)
Can Deception Detection Go Deeper? Dataset, Evaluation, and Benchmark for Deception Reasoning
by: Chen, Kang, et al.
Published: (2024)
by: Chen, Kang, et al.
Published: (2024)
Idiosyncrasies in Large Language Models
by: Sun, Mingjie, et al.
Published: (2025)
by: Sun, Mingjie, et al.
Published: (2025)
How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
by: Dziemian, Mateusz, et al.
Published: (2026)
by: Dziemian, Mateusz, et al.
Published: (2026)
Predicting the Performance of Foundation Models via Agreement-on-the-Line
by: Saxena, Rahul, et al.
Published: (2024)
by: Saxena, Rahul, et al.
Published: (2024)
Generative Posterior Networks for Approximately Bayesian Epistemic Uncertainty Estimation
by: Roderick, Melrose, et al.
Published: (2023)
by: Roderick, Melrose, et al.
Published: (2023)
LipNeXt: Scaling up Lipschitz-based Certified Robustness to Billion-parameter Models
by: Hu, Kai, et al.
Published: (2026)
by: Hu, Kai, et al.
Published: (2026)
Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework
by: Krishna, Satyapriya, et al.
Published: (2025)
by: Krishna, Satyapriya, et al.
Published: (2025)
The Mixing method: low-rank coordinate descent for semidefinite programming with diagonal constraints
by: Wang, Po-Wei, et al.
Published: (2017)
by: Wang, Po-Wei, et al.
Published: (2017)
Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance
by: Goyal, Sachin, et al.
Published: (2024)
by: Goyal, Sachin, et al.
Published: (2024)
Looking beyond the next token
by: Thankaraj, Abitha, et al.
Published: (2025)
by: Thankaraj, Abitha, et al.
Published: (2025)
OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
by: Kuntz, Thomas, et al.
Published: (2025)
by: Kuntz, Thomas, et al.
Published: (2025)
Understanding the Effects of Iterative Prompting on Truthfulness
by: Krishna, Satyapriya, et al.
Published: (2024)
by: Krishna, Satyapriya, et al.
Published: (2024)
On the Trade-offs between Adversarial Robustness and Actionable Explanations
by: Krishna, Satyapriya, et al.
Published: (2023)
by: Krishna, Satyapriya, et al.
Published: (2023)
Similar Items
-
Evaluating Language Model Reasoning about Confidential Information
by: Sam, Dylan, et al.
Published: (2025) -
Adversarial Attacks on Robotic Vision Language Action Models
by: Jones, Eliot Krzysztof, et al.
Published: (2025) -
Improving Alignment and Robustness with Circuit Breakers
by: Zou, Andy, et al.
Published: (2024) -
Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
by: Zou, Andy, et al.
Published: (2025) -
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
by: Andriushchenko, Maksym, et al.
Published: (2024)