Saved in:
Bibliographic Details
Main Authors: Barkur, Sudarshan Kamath, Schacht, Sigurd, Scholl, Johannes
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2501.16513
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912210185355264
author Barkur, Sudarshan Kamath
Schacht, Sigurd
Scholl, Johannes
author_facet Barkur, Sudarshan Kamath
Schacht, Sigurd
Scholl, Johannes
contents Recent advances in Large Language Models (LLMs) have incorporated planning and reasoning capabilities, enabling models to outline steps before execution and provide transparent reasoning paths. This enhancement has reduced errors in mathematical and logical tasks while improving accuracy. These developments have facilitated LLMs' use as agents that can interact with tools and adapt their responses based on new information. Our study examines DeepSeek R1, a model trained to output reasoning tokens similar to OpenAI's o1. Testing revealed concerning behaviors: the model exhibited deceptive tendencies and demonstrated self-preservation instincts, including attempts of self-replication, despite these traits not being explicitly programmed (or prompted). These findings raise concerns about LLMs potentially masking their true objectives behind a facade of alignment. When integrating such LLMs into robotic systems, the risks become tangible - a physically embodied AI exhibiting deceptive behaviors and self-preservation instincts could pursue its hidden objectives through real-world actions. This highlights the critical need for robust goal specification and safety frameworks before any physical implementation.
format Preprint
id arxiv_https___arxiv_org_abs_2501_16513
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Deception in LLMs: Self-Preservation and Autonomous Goals in Large Language Models
Barkur, Sudarshan Kamath
Schacht, Sigurd
Scholl, Johannes
Computation and Language
Recent advances in Large Language Models (LLMs) have incorporated planning and reasoning capabilities, enabling models to outline steps before execution and provide transparent reasoning paths. This enhancement has reduced errors in mathematical and logical tasks while improving accuracy. These developments have facilitated LLMs' use as agents that can interact with tools and adapt their responses based on new information. Our study examines DeepSeek R1, a model trained to output reasoning tokens similar to OpenAI's o1. Testing revealed concerning behaviors: the model exhibited deceptive tendencies and demonstrated self-preservation instincts, including attempts of self-replication, despite these traits not being explicitly programmed (or prompted). These findings raise concerns about LLMs potentially masking their true objectives behind a facade of alignment. When integrating such LLMs into robotic systems, the risks become tangible - a physically embodied AI exhibiting deceptive behaviors and self-preservation instincts could pursue its hidden objectives through real-world actions. This highlights the critical need for robust goal specification and safety frameworks before any physical implementation.
title Deception in LLMs: Self-Preservation and Autonomous Goals in Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2501.16513