Identifying Cooperative Personalities in Multi-agent Contexts through Personality Steering with Representation Engineering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ong, Kenneth J. K., Jun, Lye Jia, Nguyen, Hieu Minh "Jord", Cho, Seong Hah, Antolín, Natalia Pérez-Campanero
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916653876379648
author Ong, Kenneth J. K.
Jun, Lye Jia
Nguyen, Hieu Minh "Jord"
Cho, Seong Hah
Antolín, Natalia Pérez-Campanero
author_facet Ong, Kenneth J. K.
Jun, Lye Jia
Nguyen, Hieu Minh "Jord"
Cho, Seong Hah
Antolín, Natalia Pérez-Campanero
contents As Large Language Models (LLMs) gain autonomous capabilities, their coordination in multi-agent settings becomes increasingly important. However, they often struggle with cooperation, leading to suboptimal outcomes. Inspired by Axelrod's Iterated Prisoner's Dilemma (IPD) tournaments, we explore how personality traits influence LLM cooperation. Using representation engineering, we steer Big Five traits (e.g., Agreeableness, Conscientiousness) in LLMs and analyze their impact on IPD decision-making. Our results show that higher Agreeableness and Conscientiousness improve cooperation but increase susceptibility to exploitation, highlighting both the potential and limitations of personality-based steering for aligning AI agents.
format Preprint
id arxiv_https___arxiv_org_abs_2503_12722
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Identifying Cooperative Personalities in Multi-agent Contexts through Personality Steering with Representation Engineering
Ong, Kenneth J. K.
Jun, Lye Jia
Nguyen, Hieu Minh "Jord"
Cho, Seong Hah
Antolín, Natalia Pérez-Campanero
Artificial Intelligence
Computation and Language
Computer Science and Game Theory
Multiagent Systems
As Large Language Models (LLMs) gain autonomous capabilities, their coordination in multi-agent settings becomes increasingly important. However, they often struggle with cooperation, leading to suboptimal outcomes. Inspired by Axelrod's Iterated Prisoner's Dilemma (IPD) tournaments, we explore how personality traits influence LLM cooperation. Using representation engineering, we steer Big Five traits (e.g., Agreeableness, Conscientiousness) in LLMs and analyze their impact on IPD decision-making. Our results show that higher Agreeableness and Conscientiousness improve cooperation but increase susceptibility to exploitation, highlighting both the potential and limitations of personality-based steering for aligning AI agents.
title Identifying Cooperative Personalities in Multi-agent Contexts through Personality Steering with Representation Engineering
topic Artificial Intelligence
Computation and Language
Computer Science and Game Theory
Multiagent Systems
url https://arxiv.org/abs/2503.12722