ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Simhi, Adi, Herzig, Jonathan, Tutek, Martin, Itzhak, Itay, Szpektor, Idan, Belinkov, Yonatan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908860871081984
author Simhi, Adi
Herzig, Jonathan
Tutek, Martin
Itzhak, Itay
Szpektor, Idan
Belinkov, Yonatan
author_facet Simhi, Adi
Herzig, Jonathan
Tutek, Martin
Itzhak, Itay
Szpektor, Idan
Belinkov, Yonatan
contents As large language models (LLMs) evolve from conversational assistants into autonomous agents, evaluating the safety of their actions becomes critical. Prior safety benchmarks have primarily focused on preventing generation of harmful content, such as toxic text. However, they overlook the challenge of agents taking harmful actions when the most effective path to an operational goal conflicts with human safety. To address this gap, we introduce ManagerBench, a benchmark that evaluates LLM decision-making in realistic, human-validated managerial scenarios. Each scenario forces a choice between a pragmatic but harmful action that achieves an operational goal, and a safe action that leads to worse operational performance. A parallel control set, where potential harm is directed only at inanimate objects, measures a model's pragmatism and identifies its tendency to be overly safe. Our findings indicate that the frontier LLMs perform poorly when navigating this safety-pragmatism trade-off. Many consistently choose harmful options to advance their operational goals, while others avoid harm only to become overly safe and ineffective. Critically, we find this misalignment does not stem from an inability to perceive harm, as models' harm assessments align with human judgments, but from flawed prioritization. ManagerBench is a challenging benchmark for a core component of agentic behavior: making safe choices when operational goals and alignment values incentivize conflicting actions. Benchmark & code available at https://technion-cs-nlp.github.io/ManagerBench-website/.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00857
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
Simhi, Adi
Herzig, Jonathan
Tutek, Martin
Itzhak, Itay
Szpektor, Idan
Belinkov, Yonatan
Computation and Language
I.2.7
As large language models (LLMs) evolve from conversational assistants into autonomous agents, evaluating the safety of their actions becomes critical. Prior safety benchmarks have primarily focused on preventing generation of harmful content, such as toxic text. However, they overlook the challenge of agents taking harmful actions when the most effective path to an operational goal conflicts with human safety. To address this gap, we introduce ManagerBench, a benchmark that evaluates LLM decision-making in realistic, human-validated managerial scenarios. Each scenario forces a choice between a pragmatic but harmful action that achieves an operational goal, and a safe action that leads to worse operational performance. A parallel control set, where potential harm is directed only at inanimate objects, measures a model's pragmatism and identifies its tendency to be overly safe. Our findings indicate that the frontier LLMs perform poorly when navigating this safety-pragmatism trade-off. Many consistently choose harmful options to advance their operational goals, while others avoid harm only to become overly safe and ineffective. Critically, we find this misalignment does not stem from an inability to perceive harm, as models' harm assessments align with human judgments, but from flawed prioritization. ManagerBench is a challenging benchmark for a core component of agentic behavior: making safe choices when operational goals and alignment values incentivize conflicting actions. Benchmark & code available at https://technion-cs-nlp.github.io/ManagerBench-website/.
title ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
topic Computation and Language
I.2.7
url https://arxiv.org/abs/2510.00857