From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law
Fuente:
arXiv
Saved in:
| Main Authors: | Mavi, John, Găitan, Diana Teodora, Coronado, Sergio |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job?
by: Mavi, John, et al.
Published: (2024)
by: Mavi, John, et al.
Published: (2024)
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
by: Yuan, Yuan, et al.
Published: (2025)
by: Yuan, Yuan, et al.
Published: (2025)
RogueGPT: dis-ethical tuning transforms ChatGPT4 into a Rogue AI in 158 Words
by: Buscemi, Alessio, et al.
Published: (2024)
by: Buscemi, Alessio, et al.
Published: (2024)
Towards Safe Multilingual Frontier AI
by: Kanepajs, Artūrs, et al.
Published: (2024)
by: Kanepajs, Artūrs, et al.
Published: (2024)
Safe in the Future, Dangerous in the Past: Dissecting Temporal and Linguistic Vulnerabilities in LLMs
by: Said, Muhammad Abdullahi, et al.
Published: (2025)
by: Said, Muhammad Abdullahi, et al.
Published: (2025)
Cognitive Agent Compilation for Explicit Problem Solver Modeling
by: Moon, Hyeongdon, et al.
Published: (2026)
by: Moon, Hyeongdon, et al.
Published: (2026)
Knowledge Acquisition on Mass-shooting Events via LLMs for AI-Driven Justice
by: Ihugba, Benign John, et al.
Published: (2025)
by: Ihugba, Benign John, et al.
Published: (2025)
Aligning Large Language Models with Healthcare Stakeholders: A Pathway to Trustworthy AI Integration
by: Ding, Kexin, et al.
Published: (2025)
by: Ding, Kexin, et al.
Published: (2025)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
Literary Narrative as Moral Probe : A Cross-System Framework for Evaluating AI Ethical Reasoning and Refusal Behavior
by: Flynn, David C.
Published: (2026)
by: Flynn, David C.
Published: (2026)
Gender Bias in LLMs: Preliminary Evidence from Shared Parenting Scenario in Czech Family Law
by: Harasta, Jakub, et al.
Published: (2026)
by: Harasta, Jakub, et al.
Published: (2026)
Punctuated Equilibria in Artificial Intelligence: The Institutional Scaling Law and the Speciation of Sovereign AI
by: Baciak, Mark, et al.
Published: (2026)
by: Baciak, Mark, et al.
Published: (2026)
Lived Experience Not Found: LLMs Struggle to Align with Experts on Addressing Adverse Drug Reactions from Psychiatric Medication Use
by: Chandra, Mohit, et al.
Published: (2024)
by: Chandra, Mohit, et al.
Published: (2024)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
by: Choi, Sooyung, et al.
Published: (2025)
by: Choi, Sooyung, et al.
Published: (2025)
The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety
by: Rios-Sialer, Ian
Published: (2026)
by: Rios-Sialer, Ian
Published: (2026)
Standardizing Intelligence: Aligning Generative AI for Regulatory and Operational Compliance
by: Imperial, Joseph Marvin, et al.
Published: (2025)
by: Imperial, Joseph Marvin, et al.
Published: (2025)
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
by: Oh, Gyutaek, et al.
Published: (2025)
by: Oh, Gyutaek, et al.
Published: (2025)
Towards Lawful Autonomous Driving: Deriving Scenario-Aware Driving Requirements from Traffic Laws and Regulations
by: Jian, Bowen, et al.
Published: (2026)
by: Jian, Bowen, et al.
Published: (2026)
Linearly Decoding Refused Knowledge in Aligned Language Models
by: Shrivastava, Aryan, et al.
Published: (2025)
by: Shrivastava, Aryan, et al.
Published: (2025)
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Guided Persona-based AI Surveys: Can we replicate personal mobility preferences at scale using LLMs?
by: Tzachristas, Ioannis, et al.
Published: (2025)
by: Tzachristas, Ioannis, et al.
Published: (2025)
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
by: Sun, Lihao, et al.
Published: (2025)
by: Sun, Lihao, et al.
Published: (2025)
TherapyGym: Evaluating and Aligning Clinical Fidelity and Safety in Therapy Chatbots
by: Huang, Fangrui, et al.
Published: (2026)
by: Huang, Fangrui, et al.
Published: (2026)
Simulating Students with Large Language Models: A Review of Architecture, Mechanisms, and Role Modelling in Education with Generative AI
by: Marquez-Carpintero, Luis, et al.
Published: (2025)
by: Marquez-Carpintero, Luis, et al.
Published: (2025)
AI Act and Large Language Models (LLMs): When critical issues and privacy impact require human and ethical oversight
by: Fabiano, Nicola
Published: (2024)
by: Fabiano, Nicola
Published: (2024)
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
by: Yang, Chao, et al.
Published: (2024)
by: Yang, Chao, et al.
Published: (2024)
From Feature-Based Models to Generative AI: Validity Evidence for Constructed Response Scoring
by: Casabianca, Jodi M., et al.
Published: (2026)
by: Casabianca, Jodi M., et al.
Published: (2026)
From Complexity to Clarity: How AI Enhances Perceptions of Scientists and the Public's Understanding of Science
by: Markowitz, David M.
Published: (2024)
by: Markowitz, David M.
Published: (2024)
Oyster-I: Beyond Refusal -- Constructive Safety Alignment for Responsible Language Models
by: Duan, Ranjie, et al.
Published: (2025)
by: Duan, Ranjie, et al.
Published: (2025)
Giving AI Personalities Leads to More Human-Like Reasoning
by: Nighojkar, Animesh, et al.
Published: (2025)
by: Nighojkar, Animesh, et al.
Published: (2025)
The simulation of judgment in LLMs
by: Loru, Edoardo, et al.
Published: (2025)
by: Loru, Edoardo, et al.
Published: (2025)
Measuring Teaching with LLMs
by: Hardy, Michael
Published: (2025)
by: Hardy, Michael
Published: (2025)
The Political Preferences of LLMs
by: Rozado, David
Published: (2024)
by: Rozado, David
Published: (2024)
Topic Classification of Case Law Using a Large Language Model and a New Taxonomy for UK Law: AI Insights into Summary Judgment
by: Sargeant, Holli, et al.
Published: (2024)
by: Sargeant, Holli, et al.
Published: (2024)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
by: Si, Shengyun, et al.
Published: (2025)
by: Si, Shengyun, et al.
Published: (2025)
Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs
by: de Landa, Joseba Fernandez, et al.
Published: (2026)
by: de Landa, Joseba Fernandez, et al.
Published: (2026)
From Black-Box Confidence to Measurable Trust in Clinical AI: A Framework for Evidence, Supervision, and Staged Autonomy
by: Zabolotnii, Serhii, et al.
Published: (2026)
by: Zabolotnii, Serhii, et al.
Published: (2026)
SomaliBench Eval: Measuring English-to-Somali Refusal Gaps in Open-Weight Language Models
by: Dahir, Khalid Yusuf
Published: (2026)
by: Dahir, Khalid Yusuf
Published: (2026)
Moral Mazes in the Era of LLMs
by: Nguyen, Dang, et al.
Published: (2026)
by: Nguyen, Dang, et al.
Published: (2026)
SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-Ahead Reasoning
by: Wang, Lichao, et al.
Published: (2026)
by: Wang, Lichao, et al.
Published: (2026)
Similar Items
-
Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job?
by: Mavi, John, et al.
Published: (2024) -
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
by: Yuan, Yuan, et al.
Published: (2025) -
RogueGPT: dis-ethical tuning transforms ChatGPT4 into a Rogue AI in 158 Words
by: Buscemi, Alessio, et al.
Published: (2024) -
Towards Safe Multilingual Frontier AI
by: Kanepajs, Artūrs, et al.
Published: (2024) -
Safe in the Future, Dangerous in the Past: Dissecting Temporal and Linguistic Vulnerabilities in LLMs
by: Said, Muhammad Abdullahi, et al.
Published: (2025)