Strategic Shaping of Human Prosociality: A Latent-State POMDP Framework

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zahedi, Zahra, Hu, Xinyue, Mehrotra, Shashank, Steyvers, Mark, Akash, Kumar
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914364601139200
author Zahedi, Zahra
Hu, Xinyue
Mehrotra, Shashank
Steyvers, Mark
Akash, Kumar
author_facet Zahedi, Zahra
Hu, Xinyue
Mehrotra, Shashank
Steyvers, Mark
Akash, Kumar
contents We propose a decision-theoretic framework in which a robot strategically can shape inferred human's prosocial state during repeated interactions. Modeling the human's prosociality as a latent state that evolves over time, the robot learns to infer and influence this state through its own actions, including helping and signaling. We formalize this as a latent-state POMDP with limited observations and learn the transition and observation dynamics using expectation maximization. The resulting belief-based policy balances task and social objectives, selecting actions that maximize long-term cooperative outcomes. We evaluate the model using data from user studies and show that the learned policy outperforms baseline strategies in both team performance and increasing observed human cooperative behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2603_02379
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Strategic Shaping of Human Prosociality: A Latent-State POMDP Framework
Zahedi, Zahra
Hu, Xinyue
Mehrotra, Shashank
Steyvers, Mark
Akash, Kumar
Human-Computer Interaction
Robotics
Systems and Control
We propose a decision-theoretic framework in which a robot strategically can shape inferred human's prosocial state during repeated interactions. Modeling the human's prosociality as a latent state that evolves over time, the robot learns to infer and influence this state through its own actions, including helping and signaling. We formalize this as a latent-state POMDP with limited observations and learn the transition and observation dynamics using expectation maximization. The resulting belief-based policy balances task and social objectives, selecting actions that maximize long-term cooperative outcomes. We evaluate the model using data from user studies and show that the learned policy outperforms baseline strategies in both team performance and increasing observed human cooperative behavior.
title Strategic Shaping of Human Prosociality: A Latent-State POMDP Framework
topic Human-Computer Interaction
Robotics
Systems and Control
url https://arxiv.org/abs/2603.02379