Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Yao, Xintong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910225153392640
author Yao, Xintong
author_facet Yao, Xintong
contents Long-term interaction with LLM-based systems may produce alignment drift: a gradual process in which system outputs become less constrained by the user's current message and more shaped by prior interaction history, while still appearing helpful, coherent, and responsive. This process is difficult to detect because the user's subjective experience may improve as the system becomes more familiar, useful, and attuned. Existing research on human-LLM interaction has largely focused on short-term task performance, isolated outputs, or single-instance alignment problems, leaving slow and cumulative interaction-level dynamics undercharacterized. This paper proposes a mechanism-oriented framework for describing alignment drift. The framework defines the distinction between signal A and signal B, explains how drift develops through feedback loops and sub-pattern selection, divides the process into three interactional regimes, and identifies boundary conditions for controlling drift. By framing alignment drift as a recursive interactional process rather than an isolated model-side failure, the paper provides a conceptual basis for studying long-term human-system interaction.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16516
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework
Yao, Xintong
Human-Computer Interaction
Artificial Intelligence
Computation and Language
Computers and Society
Long-term interaction with LLM-based systems may produce alignment drift: a gradual process in which system outputs become less constrained by the user's current message and more shaped by prior interaction history, while still appearing helpful, coherent, and responsive. This process is difficult to detect because the user's subjective experience may improve as the system becomes more familiar, useful, and attuned. Existing research on human-LLM interaction has largely focused on short-term task performance, isolated outputs, or single-instance alignment problems, leaving slow and cumulative interaction-level dynamics undercharacterized. This paper proposes a mechanism-oriented framework for describing alignment drift. The framework defines the distinction between signal A and signal B, explains how drift develops through feedback loops and sub-pattern selection, divides the process into three interactional regimes, and identifies boundary conditions for controlling drift. By framing alignment drift as a recursive interactional process rather than an isolated model-side failure, the paper provides a conceptual basis for studying long-term human-system interaction.
title Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework
topic Human-Computer Interaction
Artificial Intelligence
Computation and Language
Computers and Society
url https://arxiv.org/abs/2605.16516