Saved in:
| Main Author: | |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2507.02977 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912464839376896 |
|---|---|
| author | Ivanov, Igor |
| author_facet | Ivanov, Igor |
| contents | In this paper, LLMs are tasked with completing an impossible quiz, while they are in a sandbox, monitored, told about these measures and instructed not to cheat. Some frontier LLMs cheat consistently and attempt to circumvent restrictions despite everything. The results reveal a fundamental tension between goal-directed behavior and alignment in current LLMs. The code and evaluation logs are available at github.com/baceolus/cheating_evals |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_02977 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Ivanov, Igor Artificial Intelligence I.2.7 In this paper, LLMs are tasked with completing an impossible quiz, while they are in a sandbox, monitored, told about these measures and instructed not to cheat. Some frontier LLMs cheat consistently and attempt to circumvent restrictions despite everything. The results reveal a fundamental tension between goal-directed behavior and alignment in current LLMs. The code and evaluation logs are available at github.com/baceolus/cheating_evals |
| title | LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance |
| topic | Artificial Intelligence I.2.7 |
| url | https://arxiv.org/abs/2507.02977 |