sudoLLM: On Multi-role Alignment of Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Saha, Soumadeep, Chaturvedi, Akshay, Mahapatra, Joy, Garain, Utpal
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909940773289984
author Saha, Soumadeep
Chaturvedi, Akshay
Mahapatra, Joy
Garain, Utpal
author_facet Saha, Soumadeep
Chaturvedi, Akshay
Mahapatra, Joy
Garain, Utpal
contents User authorization-based access privileges are a key feature in many safety-critical systems, but have not been extensively studied in the large language model (LLM) realm. In this work, drawing inspiration from such access control systems, we introduce sudoLLM, a novel framework that results in multi-role aligned LLMs, i.e., LLMs that account for, and behave in accordance with, user access rights. sudoLLM injects subtle user-based biases into queries and trains an LLM to utilize this bias signal in order to produce sensitive information if and only if the user is authorized. We present empirical results demonstrating that this approach shows substantially improved alignment, generalization, resistance to prefix-based jailbreaking attacks, and ``fails-closed''. The persistent tension between the language modeling objective and safety alignment, which is often exploited to jailbreak LLMs, is somewhat resolved with the aid of the injected bias signal. Our framework is meant as an additional security layer, and complements existing guardrail mechanisms for enhanced end-to-end safety with LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14607
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle sudoLLM: On Multi-role Alignment of Language Models
Saha, Soumadeep
Chaturvedi, Akshay
Mahapatra, Joy
Garain, Utpal
Computation and Language
Cryptography and Security
I.2.7
User authorization-based access privileges are a key feature in many safety-critical systems, but have not been extensively studied in the large language model (LLM) realm. In this work, drawing inspiration from such access control systems, we introduce sudoLLM, a novel framework that results in multi-role aligned LLMs, i.e., LLMs that account for, and behave in accordance with, user access rights. sudoLLM injects subtle user-based biases into queries and trains an LLM to utilize this bias signal in order to produce sensitive information if and only if the user is authorized. We present empirical results demonstrating that this approach shows substantially improved alignment, generalization, resistance to prefix-based jailbreaking attacks, and ``fails-closed''. The persistent tension between the language modeling objective and safety alignment, which is often exploited to jailbreak LLMs, is somewhat resolved with the aid of the injected bias signal. Our framework is meant as an additional security layer, and complements existing guardrail mechanisms for enhanced end-to-end safety with LLMs.
title sudoLLM: On Multi-role Alignment of Language Models
topic Computation and Language
Cryptography and Security
I.2.7
url https://arxiv.org/abs/2505.14607