HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Chance Jiajie, Mo, Zhenze, Tang, Yuhan, Qu, Ao, Wu, Jiayi, Zhao, Kaiya Ivy, Gan, Yulu, Fan, Jie, Yu, Jiangbo, Jiang, Hang, Liang, Paul Pu, Zhao, Jinhua, Pastor, Luis Alberto Alonso, Larson, Kent
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908634960625664
author Li, Chance Jiajie
Mo, Zhenze
Tang, Yuhan
Qu, Ao
Wu, Jiayi
Zhao, Kaiya Ivy
Gan, Yulu
Fan, Jie
Yu, Jiangbo
Jiang, Hang
Liang, Paul Pu
Zhao, Jinhua
Pastor, Luis Alberto Alonso
Larson, Kent
author_facet Li, Chance Jiajie
Mo, Zhenze
Tang, Yuhan
Qu, Ao
Wu, Jiayi
Zhao, Kaiya Ivy
Gan, Yulu
Fan, Jie
Yu, Jiangbo
Jiang, Hang
Liang, Paul Pu
Zhao, Jinhua
Pastor, Luis Alberto Alonso
Larson, Kent
contents Simulating human reasoning in open-ended tasks has long been a central aspiration in AI and cognitive science. While large language models now approximate human responses at scale, they remain tuned to population-level consensus, often erasing the individuality of reasoning styles and belief trajectories. To advance the vision of more human-like reasoning in machines, we introduce HugAgent (Human-Grounded Agent Benchmark), which rethinks human reasoning simulation along three dimensions: (i) from averaged to individualized reasoning, (ii) from behavioral mimicry to cognitive alignment, and (iii) from vignette-based to open-ended data. The benchmark evaluates whether a model can predict a specific person's behavioral responses and the underlying reasoning dynamics in out-of-distribution scenarios, given partial evidence of their prior views. HugAgent adopts a dual-track design: a human track that automates and scales the think-aloud method to collect ecologically valid human reasoning data, and a synthetic track for further scalability and systematic stress testing. This architecture enables low-cost, extensible expansion to new tasks and populations. Experiments with state-of-the-art language models reveal persistent adaptation gaps, positioning HugAgent as the first extensible benchmark for aligning machine reasoning with the individuality of human thought. The benchmark, along with its complete data collection pipeline and companion chatbot, is open-sourced as HugAgent (https://anonymous.4open.science/r/HugAgent) and TraceYourThinking (https://anonymous.4open.science/r/trace-your-thinking).
format Preprint
id arxiv_https___arxiv_org_abs_2510_15144
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning
Li, Chance Jiajie
Mo, Zhenze
Tang, Yuhan
Qu, Ao
Wu, Jiayi
Zhao, Kaiya Ivy
Gan, Yulu
Fan, Jie
Yu, Jiangbo
Jiang, Hang
Liang, Paul Pu
Zhao, Jinhua
Pastor, Luis Alberto Alonso
Larson, Kent
Artificial Intelligence
Computation and Language
Computers and Society
Simulating human reasoning in open-ended tasks has long been a central aspiration in AI and cognitive science. While large language models now approximate human responses at scale, they remain tuned to population-level consensus, often erasing the individuality of reasoning styles and belief trajectories. To advance the vision of more human-like reasoning in machines, we introduce HugAgent (Human-Grounded Agent Benchmark), which rethinks human reasoning simulation along three dimensions: (i) from averaged to individualized reasoning, (ii) from behavioral mimicry to cognitive alignment, and (iii) from vignette-based to open-ended data. The benchmark evaluates whether a model can predict a specific person's behavioral responses and the underlying reasoning dynamics in out-of-distribution scenarios, given partial evidence of their prior views. HugAgent adopts a dual-track design: a human track that automates and scales the think-aloud method to collect ecologically valid human reasoning data, and a synthetic track for further scalability and systematic stress testing. This architecture enables low-cost, extensible expansion to new tasks and populations. Experiments with state-of-the-art language models reveal persistent adaptation gaps, positioning HugAgent as the first extensible benchmark for aligning machine reasoning with the individuality of human thought. The benchmark, along with its complete data collection pipeline and companion chatbot, is open-sourced as HugAgent (https://anonymous.4open.science/r/HugAgent) and TraceYourThinking (https://anonymous.4open.science/r/trace-your-thinking).
title HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning
topic Artificial Intelligence
Computation and Language
Computers and Society
url https://arxiv.org/abs/2510.15144