Saved in:
Bibliographic Details
Main Authors: Jensen, Benjamin, Reynolds, Ian, Atalan, Yasir, Garcia, Michael, Woo, Austin, Chen, Anthony, Howarth, Trevor
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2503.06263
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910865213620224
author Jensen, Benjamin
Reynolds, Ian
Atalan, Yasir
Garcia, Michael
Woo, Austin
Chen, Anthony
Howarth, Trevor
author_facet Jensen, Benjamin
Reynolds, Ian
Atalan, Yasir
Garcia, Michael
Woo, Austin
Chen, Anthony
Howarth, Trevor
contents As national security institutions increasingly integrate Artificial Intelligence (AI) into decision-making and content generation processes, understanding the inherent biases of large language models (LLMs) is crucial. This study presents a novel benchmark designed to evaluate the biases and preferences of seven prominent foundation models-Llama 3.1 8B Instruct, Llama 3.1 70B Instruct, GPT-4o, Gemini 1.5 Pro-002, Mixtral 8x22B, Claude 3.5 Sonnet, and Qwen2 72B-in the context of international relations (IR). We designed a bias discovery study around core topics in IR using 400-expert crafted scenarios to analyze results from our selected models. These scenarios focused on four topical domains including: military escalation, military and humanitarian intervention, cooperative behavior in the international system, and alliance dynamics. Our analysis reveals noteworthy variation among model recommendations based on scenarios designed for the four tested domains. Particularly, Qwen2 72B, Gemini 1.5 Pro-002 and Llama 3.1 8B Instruct models offered significantly more escalatory recommendations than Claude 3.5 Sonnet and GPT-4o models. All models exhibit some degree of country-specific biases, often recommending less escalatory and interventionist actions for China and Russia compared to the United States and the United Kingdom. These findings highlight the necessity for controlled deployment of LLMs in high-stakes environments, emphasizing the need for domain-specific evaluations and model fine-tuning to align with institutional objectives.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06263
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Critical Foreign Policy Decisions (CFPD)-Benchmark: Measuring Diplomatic Preferences in Large Language Models
Jensen, Benjamin
Reynolds, Ian
Atalan, Yasir
Garcia, Michael
Woo, Austin
Chen, Anthony
Howarth, Trevor
Computers and Society
Artificial Intelligence
Computation and Language
Machine Learning
I.2.7
As national security institutions increasingly integrate Artificial Intelligence (AI) into decision-making and content generation processes, understanding the inherent biases of large language models (LLMs) is crucial. This study presents a novel benchmark designed to evaluate the biases and preferences of seven prominent foundation models-Llama 3.1 8B Instruct, Llama 3.1 70B Instruct, GPT-4o, Gemini 1.5 Pro-002, Mixtral 8x22B, Claude 3.5 Sonnet, and Qwen2 72B-in the context of international relations (IR). We designed a bias discovery study around core topics in IR using 400-expert crafted scenarios to analyze results from our selected models. These scenarios focused on four topical domains including: military escalation, military and humanitarian intervention, cooperative behavior in the international system, and alliance dynamics. Our analysis reveals noteworthy variation among model recommendations based on scenarios designed for the four tested domains. Particularly, Qwen2 72B, Gemini 1.5 Pro-002 and Llama 3.1 8B Instruct models offered significantly more escalatory recommendations than Claude 3.5 Sonnet and GPT-4o models. All models exhibit some degree of country-specific biases, often recommending less escalatory and interventionist actions for China and Russia compared to the United States and the United Kingdom. These findings highlight the necessity for controlled deployment of LLMs in high-stakes environments, emphasizing the need for domain-specific evaluations and model fine-tuning to align with institutional objectives.
title Critical Foreign Policy Decisions (CFPD)-Benchmark: Measuring Diplomatic Preferences in Large Language Models
topic Computers and Society
Artificial Intelligence
Computation and Language
Machine Learning
I.2.7
url https://arxiv.org/abs/2503.06263