GreedLlama: Performance of Financial Value-Aligned Large Language Models in Moral Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Jeffy, Huber, Maximilian, Tang, Kevin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914740522975232
author Yu, Jeffy
Huber, Maximilian
Tang, Kevin
author_facet Yu, Jeffy
Huber, Maximilian
Tang, Kevin
contents This paper investigates the ethical implications of aligning Large Language Models (LLMs) with financial optimization, through the case study of GreedLlama, a model fine-tuned to prioritize economically beneficial outcomes. By comparing GreedLlama's performance in moral reasoning tasks to a base Llama2 model, our results highlight a concerning trend: GreedLlama demonstrates a marked preference for profit over ethical considerations, making morally appropriate decisions at significantly lower rates than the base model in scenarios of both low and high moral ambiguity. In low ambiguity situations, GreedLlama's ethical decisions decreased to 54.4%, compared to the base model's 86.9%, while in high ambiguity contexts, the rate was 47.4% against the base model's 65.1%. These findings emphasize the risks of single-dimensional value alignment in LLMs, underscoring the need for integrating broader ethical values into AI development to ensure decisions are not solely driven by financial incentives. The study calls for a balanced approach to LLM deployment, advocating for the incorporation of ethical considerations in models intended for business applications, particularly in light of the absence of regulatory oversight.
format Preprint
id arxiv_https___arxiv_org_abs_2404_02934
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GreedLlama: Performance of Financial Value-Aligned Large Language Models in Moral Reasoning
Yu, Jeffy
Huber, Maximilian
Tang, Kevin
Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
This paper investigates the ethical implications of aligning Large Language Models (LLMs) with financial optimization, through the case study of GreedLlama, a model fine-tuned to prioritize economically beneficial outcomes. By comparing GreedLlama's performance in moral reasoning tasks to a base Llama2 model, our results highlight a concerning trend: GreedLlama demonstrates a marked preference for profit over ethical considerations, making morally appropriate decisions at significantly lower rates than the base model in scenarios of both low and high moral ambiguity. In low ambiguity situations, GreedLlama's ethical decisions decreased to 54.4%, compared to the base model's 86.9%, while in high ambiguity contexts, the rate was 47.4% against the base model's 65.1%. These findings emphasize the risks of single-dimensional value alignment in LLMs, underscoring the need for integrating broader ethical values into AI development to ensure decisions are not solely driven by financial incentives. The study calls for a balanced approach to LLM deployment, advocating for the incorporation of ethical considerations in models intended for business applications, particularly in light of the absence of regulatory oversight.
title GreedLlama: Performance of Financial Value-Aligned Large Language Models in Moral Reasoning
topic Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
url https://arxiv.org/abs/2404.02934