Same Performance, Hidden Bias: Evaluating Hypothesis- and Recommendation-Driven AI

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Benk, Michaela, Miller, Tim
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911520870367232
author Benk, Michaela
Miller, Tim
author_facet Benk, Michaela
Miller, Tim
contents The HCI community commonly evaluates decision support systems based on whether they improve task performance or promote appropriate user reliance. In this work, we look beyond decision outcomes to examine the process through which users develop decision-making strategies. Through a web-based experiment (N = 290) comparing recommendation-driven and hypothesis-driven interaction designs, and using Signal Detection Theory as a theoretical framework, we show that even when performance remains identical, recommendation-driven designs lower participants' thresholds for sufficient evidence and introduce a "hidden bias" in their judgments, resulting in a shifted distribution of errors. Furthermore, we find that experts are just as susceptible to these systemic shifts as novices. We conclude by advocating for a shift in focus: prioritizing decision processes and the preservation of stable evidence standards over performance and reliance alone.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15824
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Same Performance, Hidden Bias: Evaluating Hypothesis- and Recommendation-Driven AI
Benk, Michaela
Miller, Tim
Human-Computer Interaction
The HCI community commonly evaluates decision support systems based on whether they improve task performance or promote appropriate user reliance. In this work, we look beyond decision outcomes to examine the process through which users develop decision-making strategies. Through a web-based experiment (N = 290) comparing recommendation-driven and hypothesis-driven interaction designs, and using Signal Detection Theory as a theoretical framework, we show that even when performance remains identical, recommendation-driven designs lower participants' thresholds for sufficient evidence and introduce a "hidden bias" in their judgments, resulting in a shifted distribution of errors. Furthermore, we find that experts are just as susceptible to these systemic shifts as novices. We conclude by advocating for a shift in focus: prioritizing decision processes and the preservation of stable evidence standards over performance and reliance alone.
title Same Performance, Hidden Bias: Evaluating Hypothesis- and Recommendation-Driven AI
topic Human-Computer Interaction
url https://arxiv.org/abs/2603.15824