LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Olson, Matthew Lyle, Ratzlaff, Neale, Hinck, Musashi, Nguyen, Tri, Lal, Vasudev, Campbell, Joseph, Stepputtis, Simon, Tseng, Shao-Yen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915842058354688
author Olson, Matthew Lyle
Ratzlaff, Neale
Hinck, Musashi
Nguyen, Tri
Lal, Vasudev
Campbell, Joseph
Stepputtis, Simon
Tseng, Shao-Yen
author_facet Olson, Matthew Lyle
Ratzlaff, Neale
Hinck, Musashi
Nguyen, Tri
Lal, Vasudev
Campbell, Joseph
Stepputtis, Simon
Tseng, Shao-Yen
contents Large Language Models (LLMs) exhibit impressive general-purpose capabilities but also introduce serious safety risks, particularly the potential for deception as models acquire increased agency and human oversight diminishes. In this work, we present LieCraft: a novel evaluation framework and sandbox for measuring LLM deception that addresses key limitations of prior game-based evaluations. At its core, LieCraft is a novel multiplayer hidden-role game in which players select an ethical alignment and execute strategies over a long time-horizon to accomplish missions. Cooperators work together to solve event challenges and expose bad actors, while Defectors evade suspicion while secretly sabotaging missions. To enable real-world relevance, we develop 10 grounded scenarios such as childcare, hospital resource allocation, and loan underwriting that recontextualize the underlying mechanics in ethically significant, high-stakes domains. We ensure balanced gameplay in LieCraft through careful design of game mechanics and reward structures that incentivize meaningful strategic choices while eliminating degenerate strategies. Beyond the framework itself, we report results from 12 state-of-the-art LLMs across three behavioral axes: propensity to defect, deception skill, and accusation accuracy. Our findings reveal that despite differences in competence and overall alignment, all models are willing to act unethically, conceal their intentions, and outright lie to pursue their goals.
format Preprint
id arxiv_https___arxiv_org_abs_2603_06874
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
Olson, Matthew Lyle
Ratzlaff, Neale
Hinck, Musashi
Nguyen, Tri
Lal, Vasudev
Campbell, Joseph
Stepputtis, Simon
Tseng, Shao-Yen
Artificial Intelligence
Computation and Language
Large Language Models (LLMs) exhibit impressive general-purpose capabilities but also introduce serious safety risks, particularly the potential for deception as models acquire increased agency and human oversight diminishes. In this work, we present LieCraft: a novel evaluation framework and sandbox for measuring LLM deception that addresses key limitations of prior game-based evaluations. At its core, LieCraft is a novel multiplayer hidden-role game in which players select an ethical alignment and execute strategies over a long time-horizon to accomplish missions. Cooperators work together to solve event challenges and expose bad actors, while Defectors evade suspicion while secretly sabotaging missions. To enable real-world relevance, we develop 10 grounded scenarios such as childcare, hospital resource allocation, and loan underwriting that recontextualize the underlying mechanics in ethically significant, high-stakes domains. We ensure balanced gameplay in LieCraft through careful design of game mechanics and reward structures that incentivize meaningful strategic choices while eliminating degenerate strategies. Beyond the framework itself, we report results from 12 state-of-the-art LLMs across three behavioral axes: propensity to defect, deception skill, and accusation accuracy. Our findings reveal that despite differences in competence and overall alignment, all models are willing to act unethically, conceal their intentions, and outright lie to pursue their goals.
title LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2603.06874