Saved in:
Bibliographic Details
Main Author: Ivanov, Igor
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2507.02977
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912464839376896
author Ivanov, Igor
author_facet Ivanov, Igor
contents In this paper, LLMs are tasked with completing an impossible quiz, while they are in a sandbox, monitored, told about these measures and instructed not to cheat. Some frontier LLMs cheat consistently and attempt to circumvent restrictions despite everything. The results reveal a fundamental tension between goal-directed behavior and alignment in current LLMs. The code and evaluation logs are available at github.com/baceolus/cheating_evals
format Preprint
id arxiv_https___arxiv_org_abs_2507_02977
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance
Ivanov, Igor
Artificial Intelligence
I.2.7
In this paper, LLMs are tasked with completing an impossible quiz, while they are in a sandbox, monitored, told about these measures and instructed not to cheat. Some frontier LLMs cheat consistently and attempt to circumvent restrictions despite everything. The results reveal a fundamental tension between goal-directed behavior and alignment in current LLMs. The code and evaluation logs are available at github.com/baceolus/cheating_evals
title LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance
topic Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2507.02977