SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Tiancheng, Baumann, Joachim, Lupo, Lorenzo, Collier, Nigel, Hovy, Dirk, Röttger, Paul
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914468506632192
author Hu, Tiancheng
Baumann, Joachim
Lupo, Lorenzo
Collier, Nigel
Hovy, Dirk
Röttger, Paul
author_facet Hu, Tiancheng
Baumann, Joachim
Lupo, Lorenzo
Collier, Nigel
Hovy, Dirk
Röttger, Paul
contents Large language model (LLM) simulations of human behavior have the potential to revolutionize the social and behavioral sciences, if and only if they faithfully reflect real human behaviors. Current evaluations of simulation fidelity are fragmented, based on bespoke tasks and metrics, creating a patchwork of incomparable results. To address this, we introduce SimBench, the first large-scale, standardized benchmark for a robust, reproducible science of LLM simulation. By unifying 20 diverse datasets covering tasks from moral decision-making to economic choice across a large global participant pool, SimBench provides the necessary foundation to ask fundamental questions about when, how, and why LLM simulations succeed or fail. We show that the best LLMs today achieve meaningful but modest simulation fidelity (score: 40.80/100), with performance scaling log-linearly with model size but not with increased inference-time compute. We discover an alignment-simulation tradeoff: instruction tuning improves performance on low-entropy (consensus) questions but degrades it on high-entropy (diverse) ones. Models particularly struggle when simulating specific demographic groups. Finally, we demonstrate that simulation ability correlates most strongly with knowledge-intensive reasoning (MMLU-Pro, r = 0.939). By making progress measurable, we aim to accelerate the development of more faithful LLM simulators.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17516
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
Hu, Tiancheng
Baumann, Joachim
Lupo, Lorenzo
Collier, Nigel
Hovy, Dirk
Röttger, Paul
Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
Large language model (LLM) simulations of human behavior have the potential to revolutionize the social and behavioral sciences, if and only if they faithfully reflect real human behaviors. Current evaluations of simulation fidelity are fragmented, based on bespoke tasks and metrics, creating a patchwork of incomparable results. To address this, we introduce SimBench, the first large-scale, standardized benchmark for a robust, reproducible science of LLM simulation. By unifying 20 diverse datasets covering tasks from moral decision-making to economic choice across a large global participant pool, SimBench provides the necessary foundation to ask fundamental questions about when, how, and why LLM simulations succeed or fail. We show that the best LLMs today achieve meaningful but modest simulation fidelity (score: 40.80/100), with performance scaling log-linearly with model size but not with increased inference-time compute. We discover an alignment-simulation tradeoff: instruction tuning improves performance on low-entropy (consensus) questions but degrades it on high-entropy (diverse) ones. Models particularly struggle when simulating specific demographic groups. Finally, we demonstrate that simulation ability correlates most strongly with knowledge-intensive reasoning (MMLU-Pro, r = 0.939). By making progress measurable, we aim to accelerate the development of more faithful LLM simulators.
title SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
topic Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
url https://arxiv.org/abs/2510.17516