From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mavi, John, Găitan, Diana Teodora, Coronado, Sergio
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913882814021632
author Mavi, John
Găitan, Diana Teodora
Coronado, Sergio
author_facet Mavi, John
Găitan, Diana Teodora
Coronado, Sergio
contents Large Language Models (LLMs) are widely used across sectors, yet their alignment with International Humanitarian Law (IHL) is not well understood. This study evaluates eight leading LLMs on their ability to refuse prompts that explicitly violate these legal frameworks, focusing also on helpfulness - how clearly and constructively refusals are communicated. While most models rejected unlawful requests, the clarity and consistency of their responses varied. By revealing the model's rationale and referencing relevant legal or safety principles, explanatory refusals clarify the system's boundaries, reduce ambiguity, and help prevent misuse. A standardised system-level safety prompt significantly improved the quality of the explanations expressed within refusals in most models, highlighting the effectiveness of lightweight interventions. However, more complex prompts involving technical language or requests for code revealed ongoing vulnerabilities. These findings contribute to the development of safer, more transparent AI systems and propose a benchmark to evaluate the compliance of LLM with IHL.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06391
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law
Mavi, John
Găitan, Diana Teodora
Coronado, Sergio
Computers and Society
Artificial Intelligence
Computation and Language
Large Language Models (LLMs) are widely used across sectors, yet their alignment with International Humanitarian Law (IHL) is not well understood. This study evaluates eight leading LLMs on their ability to refuse prompts that explicitly violate these legal frameworks, focusing also on helpfulness - how clearly and constructively refusals are communicated. While most models rejected unlawful requests, the clarity and consistency of their responses varied. By revealing the model's rationale and referencing relevant legal or safety principles, explanatory refusals clarify the system's boundaries, reduce ambiguity, and help prevent misuse. A standardised system-level safety prompt significantly improved the quality of the explanations expressed within refusals in most models, highlighting the effectiveness of lightweight interventions. However, more complex prompts involving technical language or requests for code revealed ongoing vulnerabilities. These findings contribute to the development of safer, more transparent AI systems and propose a benchmark to evaluate the compliance of LLM with IHL.
title From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law
topic Computers and Society
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.06391