Towards Safe and Honest AI Agents with Neural Self-Other Overlap

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Carauleanu, Marc, Vaiana, Michael, Rosenblatt, Judd, Berg, Cameron, de Lucena, Diogo Schwerz
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909437564813312
author Carauleanu, Marc
Vaiana, Michael
Rosenblatt, Judd
Berg, Cameron
de Lucena, Diogo Schwerz
author_facet Carauleanu, Marc
Vaiana, Michael
Rosenblatt, Judd
Berg, Cameron
de Lucena, Diogo Schwerz
contents As AI systems increasingly make critical decisions, deceptive AI poses a significant challenge to trust and safety. We present Self-Other Overlap (SOO) fine-tuning, a promising approach in AI Safety that could substantially improve our ability to build honest artificial intelligence. Inspired by cognitive neuroscience research on empathy, SOO aims to align how AI models represent themselves and others. Our experiments on LLMs with 7B, 27B, and 78B parameters demonstrate SOO's efficacy: deceptive responses of Mistral-7B-Instruct-v0.2 dropped from 73.6% to 17.2% with no observed reduction in general task performance, while in Gemma-2-27b-it and CalmeRys-78B-Orpo-v0.1 deceptive responses were reduced from 100% to 9.3% and 2.7%, respectively, with a small impact on capabilities. In reinforcement learning scenarios, SOO-trained agents showed significantly reduced deceptive behavior. SOO's focus on contrastive self and other-referencing observations offers strong potential for generalization across AI architectures. While current applications focus on language models and simple RL environments, SOO could pave the way for more trustworthy AI in broader domains. Ethical implications and long-term effects warrant further investigation, but SOO represents a significant step forward in AI safety research.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16325
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Safe and Honest AI Agents with Neural Self-Other Overlap
Carauleanu, Marc
Vaiana, Michael
Rosenblatt, Judd
Berg, Cameron
de Lucena, Diogo Schwerz
Artificial Intelligence
Cryptography and Security
As AI systems increasingly make critical decisions, deceptive AI poses a significant challenge to trust and safety. We present Self-Other Overlap (SOO) fine-tuning, a promising approach in AI Safety that could substantially improve our ability to build honest artificial intelligence. Inspired by cognitive neuroscience research on empathy, SOO aims to align how AI models represent themselves and others. Our experiments on LLMs with 7B, 27B, and 78B parameters demonstrate SOO's efficacy: deceptive responses of Mistral-7B-Instruct-v0.2 dropped from 73.6% to 17.2% with no observed reduction in general task performance, while in Gemma-2-27b-it and CalmeRys-78B-Orpo-v0.1 deceptive responses were reduced from 100% to 9.3% and 2.7%, respectively, with a small impact on capabilities. In reinforcement learning scenarios, SOO-trained agents showed significantly reduced deceptive behavior. SOO's focus on contrastive self and other-referencing observations offers strong potential for generalization across AI architectures. While current applications focus on language models and simple RL environments, SOO could pave the way for more trustworthy AI in broader domains. Ethical implications and long-term effects warrant further investigation, but SOO represents a significant step forward in AI safety research.
title Towards Safe and Honest AI Agents with Neural Self-Other Overlap
topic Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2412.16325