An Affective-Taxis Hypothesis for Alignment and Interpretability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sennesh, Eli, Ramstead, Maxwell
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909620794032128
author Sennesh, Eli
Ramstead, Maxwell
author_facet Sennesh, Eli
Ramstead, Maxwell
contents AI alignment is a field of research that aims to develop methods to ensure that agents always behave in a manner aligned with (i.e. consistently with) the goals and values of their human operators, no matter their level of capability. This paper proposes an affectivist approach to the alignment problem, re-framing the concepts of goals and values in terms of affective taxis, and explaining the emergence of affective valence by appealing to recent work in evolutionary-developmental and computational neuroscience. We review the state of the art and, building on this work, we propose a computational model of affect based on taxis navigation. We discuss evidence in a tractable model organism that our model reflects aspects of biological taxis navigation. We conclude with a discussion of the role of affective taxis in AI alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17024
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Affective-Taxis Hypothesis for Alignment and Interpretability
Sennesh, Eli
Ramstead, Maxwell
Artificial Intelligence
Neurons and Cognition
AI alignment is a field of research that aims to develop methods to ensure that agents always behave in a manner aligned with (i.e. consistently with) the goals and values of their human operators, no matter their level of capability. This paper proposes an affectivist approach to the alignment problem, re-framing the concepts of goals and values in terms of affective taxis, and explaining the emergence of affective valence by appealing to recent work in evolutionary-developmental and computational neuroscience. We review the state of the art and, building on this work, we propose a computational model of affect based on taxis navigation. We discuss evidence in a tractable model organism that our model reflects aspects of biological taxis navigation. We conclude with a discussion of the role of affective taxis in AI alignment.
title An Affective-Taxis Hypothesis for Alignment and Interpretability
topic Artificial Intelligence
Neurons and Cognition
url https://arxiv.org/abs/2505.17024