Identifying Cooperative Personalities in Multi-agent Contexts through Personality Steering with Representation Engineering
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916653876379648 |
|---|---|
| author | Ong, Kenneth J. K. Jun, Lye Jia Nguyen, Hieu Minh "Jord" Cho, Seong Hah Antolín, Natalia Pérez-Campanero |
| author_facet | Ong, Kenneth J. K. Jun, Lye Jia Nguyen, Hieu Minh "Jord" Cho, Seong Hah Antolín, Natalia Pérez-Campanero |
| contents | As Large Language Models (LLMs) gain autonomous capabilities, their coordination in multi-agent settings becomes increasingly important. However, they often struggle with cooperation, leading to suboptimal outcomes. Inspired by Axelrod's Iterated Prisoner's Dilemma (IPD) tournaments, we explore how personality traits influence LLM cooperation. Using representation engineering, we steer Big Five traits (e.g., Agreeableness, Conscientiousness) in LLMs and analyze their impact on IPD decision-making. Our results show that higher Agreeableness and Conscientiousness improve cooperation but increase susceptibility to exploitation, highlighting both the potential and limitations of personality-based steering for aligning AI agents. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_12722 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Identifying Cooperative Personalities in Multi-agent Contexts through Personality Steering with Representation Engineering Ong, Kenneth J. K. Jun, Lye Jia Nguyen, Hieu Minh "Jord" Cho, Seong Hah Antolín, Natalia Pérez-Campanero Artificial Intelligence Computation and Language Computer Science and Game Theory Multiagent Systems As Large Language Models (LLMs) gain autonomous capabilities, their coordination in multi-agent settings becomes increasingly important. However, they often struggle with cooperation, leading to suboptimal outcomes. Inspired by Axelrod's Iterated Prisoner's Dilemma (IPD) tournaments, we explore how personality traits influence LLM cooperation. Using representation engineering, we steer Big Five traits (e.g., Agreeableness, Conscientiousness) in LLMs and analyze their impact on IPD decision-making. Our results show that higher Agreeableness and Conscientiousness improve cooperation but increase susceptibility to exploitation, highlighting both the potential and limitations of personality-based steering for aligning AI agents. |
| title | Identifying Cooperative Personalities in Multi-agent Contexts through Personality Steering with Representation Engineering |
| topic | Artificial Intelligence Computation and Language Computer Science and Game Theory Multiagent Systems |
| url | https://arxiv.org/abs/2503.12722 |