Assessing Policy Updates: Toward Trust-Preserving Intelligent User Interfaces

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Solomon, Matan, Amir, Ofra, Ben-Porat, Omer
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912644144824320
author Solomon, Matan
Amir, Ofra
Ben-Porat, Omer
author_facet Solomon, Matan
Amir, Ofra
Ben-Porat, Omer
contents Reinforcement learning agents are often updated with human feedback, yet such updates can be unreliable: reward misspecification, preference conflicts, or limited data may leave policies unchanged or even worse. Because policies are difficult to interpret directly, users face the challenge of deciding whether an update has truly helped. We propose that assessing model updates -- not just a single model -- is a critical design challenge for intelligent user interfaces. In a controlled study, participants provided feedback to an agent in a gridworld and then compared its original and updated policies. We evaluated four strategies for communicating updates: no demonstration, same-context, random-context, and salient-contrast demonstrations designed to highlight informative differences. Salient-contrast demonstrations significantly improved participants' ability to detect when updates helped or harmed performance, mitigating participants' bias towards assuming that feedback is always beneficial, and supported better trust calibration across contexts.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10616
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Assessing Policy Updates: Toward Trust-Preserving Intelligent User Interfaces
Solomon, Matan
Amir, Ofra
Ben-Porat, Omer
Human-Computer Interaction
Reinforcement learning agents are often updated with human feedback, yet such updates can be unreliable: reward misspecification, preference conflicts, or limited data may leave policies unchanged or even worse. Because policies are difficult to interpret directly, users face the challenge of deciding whether an update has truly helped. We propose that assessing model updates -- not just a single model -- is a critical design challenge for intelligent user interfaces. In a controlled study, participants provided feedback to an agent in a gridworld and then compared its original and updated policies. We evaluated four strategies for communicating updates: no demonstration, same-context, random-context, and salient-contrast demonstrations designed to highlight informative differences. Salient-contrast demonstrations significantly improved participants' ability to detect when updates helped or harmed performance, mitigating participants' bias towards assuming that feedback is always beneficial, and supported better trust calibration across contexts.
title Assessing Policy Updates: Toward Trust-Preserving Intelligent User Interfaces
topic Human-Computer Interaction
url https://arxiv.org/abs/2510.10616