Frictional Agent Alignment Framework: Slow Down and Don't Break Things

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nath, Abhijnan, Graff, Carine, Bachinin, Andrei, Krishnaswamy, Nikhil
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912394588979200
author Nath, Abhijnan
Graff, Carine
Bachinin, Andrei
Krishnaswamy, Nikhil
author_facet Nath, Abhijnan
Graff, Carine
Bachinin, Andrei
Krishnaswamy, Nikhil
contents AI support of collaborative interactions entails mediating potential misalignment between interlocutor beliefs. Common preference alignment methods like DPO excel in static settings, but struggle in dynamic collaborative tasks where the explicit signals of interlocutor beliefs are sparse and skewed. We propose the Frictional Agent Alignment Framework (FAAF), to generate precise, context-aware "friction" that prompts for deliberation and re-examination of existing evidence. FAAF's two-player objective decouples from data skew: a frictive-state policy identifies belief misalignments, while an intervention policy crafts collaborator-preferred responses. We derive an analytical solution to this objective, enabling training a single policy via a simple supervised loss. Experiments on three benchmarks show FAAF outperforms competitors in producing concise, interpretable friction and in OOD generalization. By aligning LLMs to act as adaptive "thought partners" -- not passive responders -- FAAF advances scalable, dynamic human-AI collaboration. Our code and data can be found at https://github.com/csu-signal/FAAF_ACL.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19428
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Frictional Agent Alignment Framework: Slow Down and Don't Break Things
Nath, Abhijnan
Graff, Carine
Bachinin, Andrei
Krishnaswamy, Nikhil
Computation and Language
AI support of collaborative interactions entails mediating potential misalignment between interlocutor beliefs. Common preference alignment methods like DPO excel in static settings, but struggle in dynamic collaborative tasks where the explicit signals of interlocutor beliefs are sparse and skewed. We propose the Frictional Agent Alignment Framework (FAAF), to generate precise, context-aware "friction" that prompts for deliberation and re-examination of existing evidence. FAAF's two-player objective decouples from data skew: a frictive-state policy identifies belief misalignments, while an intervention policy crafts collaborator-preferred responses. We derive an analytical solution to this objective, enabling training a single policy via a simple supervised loss. Experiments on three benchmarks show FAAF outperforms competitors in producing concise, interpretable friction and in OOD generalization. By aligning LLMs to act as adaptive "thought partners" -- not passive responders -- FAAF advances scalable, dynamic human-AI collaboration. Our code and data can be found at https://github.com/csu-signal/FAAF_ACL.
title Frictional Agent Alignment Framework: Slow Down and Don't Break Things
topic Computation and Language
url https://arxiv.org/abs/2505.19428