Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Heitzig, Jobst, Potham, Ram
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913974930374656
author Heitzig, Jobst
Potham, Ram
author_facet Heitzig, Jobst
Potham, Ram
contents Power is a key concept in AI safety: power-seeking as an instrumental goal, sudden or gradual disempowerment of humans, power balance in human-AI interaction and international AI governance. At the same time, power as the ability to pursue diverse goals is essential for wellbeing. This paper explores the idea of promoting both safety and wellbeing by forcing AI agents explicitly to empower humans and to manage the power balance between humans and AI agents in a desirable way. Using a principled, partially axiomatic approach, we design a parametrizable and decomposable objective function that represents an inequality- and risk-averse long-term aggregate of human power. It takes into account humans' bounded rationality and social norms, and, crucially, considers a wide variety of possible human goals. We derive algorithms for computing that metric by backward induction or approximating it via a form of multi-agent reinforcement learning from a given world model. We exemplify the consequences of (softly) maximizing this metric in a variety of paradigmatic situations and describe what instrumental sub-goals it will likely imply. Our cautious assessment is that softly maximizing suitable aggregate metrics of human power might constitute a beneficial objective for agentic AI systems that is safer than direct utility-based objectives.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00159
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power
Heitzig, Jobst
Potham, Ram
Artificial Intelligence
Computers and Society
Machine Learning
Theoretical Economics
Optimization and Control
68Txx
I.2
Power is a key concept in AI safety: power-seeking as an instrumental goal, sudden or gradual disempowerment of humans, power balance in human-AI interaction and international AI governance. At the same time, power as the ability to pursue diverse goals is essential for wellbeing. This paper explores the idea of promoting both safety and wellbeing by forcing AI agents explicitly to empower humans and to manage the power balance between humans and AI agents in a desirable way. Using a principled, partially axiomatic approach, we design a parametrizable and decomposable objective function that represents an inequality- and risk-averse long-term aggregate of human power. It takes into account humans' bounded rationality and social norms, and, crucially, considers a wide variety of possible human goals. We derive algorithms for computing that metric by backward induction or approximating it via a form of multi-agent reinforcement learning from a given world model. We exemplify the consequences of (softly) maximizing this metric in a variety of paradigmatic situations and describe what instrumental sub-goals it will likely imply. Our cautious assessment is that softly maximizing suitable aggregate metrics of human power might constitute a beneficial objective for agentic AI systems that is safer than direct utility-based objectives.
title Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power
topic Artificial Intelligence
Computers and Society
Machine Learning
Theoretical Economics
Optimization and Control
68Txx
I.2
url https://arxiv.org/abs/2508.00159