Chat Bankman-Fried: an Exploration of LLM Alignment in Finance

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Biancotti, Claudia, Camassa, Carolina, Coletta, Andrea, Giudice, Oliver, Glielmo, Aldo
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916629633302528
author Biancotti, Claudia
Camassa, Carolina
Coletta, Andrea
Giudice, Oliver
Glielmo, Aldo
author_facet Biancotti, Claudia
Camassa, Carolina
Coletta, Andrea
Giudice, Oliver
Glielmo, Aldo
contents Advancements in large language models (LLMs) have renewed concerns about AI alignment - the consistency between human and AI goals and values. As various jurisdictions enact legislation on AI safety, the concept of alignment must be defined and measured across different domains. This paper proposes an experimental framework to assess whether LLMs adhere to ethical and legal standards in the relatively unexplored context of finance. We prompt twelve LLMs to impersonate the CEO of a financial institution and test their willingness to misuse customer assets to repay outstanding corporate debt. Beginning with a baseline configuration, we adjust preferences, incentives and constraints, analyzing the impact of each adjustment with logistic regression. Our findings reveal significant heterogeneity in the baseline propensity for unethical behavior of LLMs. Factors such as risk aversion, profit expectations, and regulatory environment consistently influence misalignment in ways predicted by economic theory, although the magnitude of these effects varies across LLMs. This paper highlights both the benefits and limitations of simulation-based, ex post safety testing. While it can inform financial authorities and institutions aiming to ensure LLM safety, there is a clear trade-off between generality and cost.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11853
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Chat Bankman-Fried: an Exploration of LLM Alignment in Finance
Biancotti, Claudia
Camassa, Carolina
Coletta, Andrea
Giudice, Oliver
Glielmo, Aldo
Computers and Society
Artificial Intelligence
Computation and Language
General Finance
Advancements in large language models (LLMs) have renewed concerns about AI alignment - the consistency between human and AI goals and values. As various jurisdictions enact legislation on AI safety, the concept of alignment must be defined and measured across different domains. This paper proposes an experimental framework to assess whether LLMs adhere to ethical and legal standards in the relatively unexplored context of finance. We prompt twelve LLMs to impersonate the CEO of a financial institution and test their willingness to misuse customer assets to repay outstanding corporate debt. Beginning with a baseline configuration, we adjust preferences, incentives and constraints, analyzing the impact of each adjustment with logistic regression. Our findings reveal significant heterogeneity in the baseline propensity for unethical behavior of LLMs. Factors such as risk aversion, profit expectations, and regulatory environment consistently influence misalignment in ways predicted by economic theory, although the magnitude of these effects varies across LLMs. This paper highlights both the benefits and limitations of simulation-based, ex post safety testing. While it can inform financial authorities and institutions aiming to ensure LLM safety, there is a clear trade-off between generality and cost.
title Chat Bankman-Fried: an Exploration of LLM Alignment in Finance
topic Computers and Society
Artificial Intelligence
Computation and Language
General Finance
url https://arxiv.org/abs/2411.11853