Benchmarking LLMs for Community Governance Simulation with Life-history Narratives
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Xu, Li, Yuanzi, Wang, Lei, Lu, Nan, Wang, Yang, Wang, Anding, Shi, Lei, Fu, Xiaoxing, Wen, Ji-Rong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
YuLan-OneSim: Towards the Next Generation of Social Simulator with Large Language Models
por: Wang, Lei, et al.
Publicado: (2025)
por: Wang, Lei, et al.
Publicado: (2025)
How Far Are LLMs from Believable AI? A Benchmark for Evaluating the Believability of Human Behavior Simulation
por: Xiao, Yang, et al.
Publicado: (2023)
por: Xiao, Yang, et al.
Publicado: (2023)
LLM Agents as Social Scientists: A Human-AI Collaborative Platform for Social Science Automation
por: Wang, Lei, et al.
Publicado: (2026)
por: Wang, Lei, et al.
Publicado: (2026)
When Assurance Undermines Intelligence: The Efficiency Costs of Data Governance in AI-Enabled Labor Markets
por: Chen, Lei, et al.
Publicado: (2025)
por: Chen, Lei, et al.
Publicado: (2025)
Benchmarking LLMs for Political Science: A United Nations Perspective
por: Liang, Yueqing, et al.
Publicado: (2025)
por: Liang, Yueqing, et al.
Publicado: (2025)
Fully Dense α‐SiC Ceramics With Enhanced Strength and Toughness Fabricated by High‐Pressure Sintering
por: Yuanpei Lei, et al.
Publicado: (2026)
por: Yuanpei Lei, et al.
Publicado: (2026)
Leveraging LLM-based agents for social science research: insights from citation network simulations
por: Ji, Jiarui, et al.
Publicado: (2025)
por: Ji, Jiarui, et al.
Publicado: (2025)
WARBENCH: A Comprehensive Benchmark for Evaluating LLMs in Military Decision-Making
por: Li, Zongjie, et al.
Publicado: (2026)
por: Li, Zongjie, et al.
Publicado: (2026)
Governable AI: Provable Safety Under Extreme Threat Models
por: Wang, Donglin, et al.
Publicado: (2025)
por: Wang, Donglin, et al.
Publicado: (2025)
ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases
por: Li, Yuchong, et al.
Publicado: (2025)
por: Li, Yuchong, et al.
Publicado: (2025)
Spark plasma sintering of sodium bismuth niobate that exhibits superior piezoelectric performance
por: Guo‐Hao Li, et al.
Publicado: (2024)
por: Guo‐Hao Li, et al.
Publicado: (2024)
Rethinking Publication: A Certification Framework for AI-Enabled Research
por: Lu, Yang, et al.
Publicado: (2026)
por: Lu, Yang, et al.
Publicado: (2026)
AppellateGen: A Benchmark for Appellate Legal Judgment Generation
por: Yang, Hongkun, et al.
Publicado: (2026)
por: Yang, Hongkun, et al.
Publicado: (2026)
From ChatGPT to DeepSeek: Can LLMs Simulate Humanity?
por: Wang, Qian, et al.
Publicado: (2025)
por: Wang, Qian, et al.
Publicado: (2025)
GGBound: A Genome-Grounded Agent for Microbial Life-Boundary Prediction
por: Huang, Hanbo, et al.
Publicado: (2026)
por: Huang, Hanbo, et al.
Publicado: (2026)
A Benchmark for Fairness-Aware Graph Learning
por: Dong, Yushun, et al.
Publicado: (2024)
por: Dong, Yushun, et al.
Publicado: (2024)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
por: Majumdar, Ayan, et al.
Publicado: (2025)
por: Majumdar, Ayan, et al.
Publicado: (2025)
Evolutionary vaccination dynamics under higher-order reinforcement pressure
por: Lu, Yikang, et al.
Publicado: (2026)
por: Lu, Yikang, et al.
Publicado: (2026)
Unbiased third-party bots lead to a tradeoff between cooperation and social payoffs
por: He, Zhixue, et al.
Publicado: (2024)
por: He, Zhixue, et al.
Publicado: (2024)
Evolutionary Dynamics of Reputation-Based Voluntary Prisoner's Dilemma Games
por: Shen, Chen, et al.
Publicado: (2026)
por: Shen, Chen, et al.
Publicado: (2026)
HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning
por: Li, Chance Jiajie, et al.
Publicado: (2025)
por: Li, Chance Jiajie, et al.
Publicado: (2025)
Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits
por: Ye, Jinyi, et al.
Publicado: (2026)
por: Ye, Jinyi, et al.
Publicado: (2026)
Simple bots breed social punishment in humans
por: Shen, Chen, et al.
Publicado: (2022)
por: Shen, Chen, et al.
Publicado: (2022)
Exit options sustain altruistic punishment and decrease the second-order free-riders, but it is not a panacea
por: Shen, Chen, et al.
Publicado: (2023)
por: Shen, Chen, et al.
Publicado: (2023)
Narrating China's Governance
por: People's Daily, Department of Commentary
Publicado: (2020)
por: People's Daily, Department of Commentary
Publicado: (2020)
How Committed Individuals Shape Social Dynamics: A Survey on Coordination Games and Social Dilemma Games
por: Shen, Chen, et al.
Publicado: (2023)
por: Shen, Chen, et al.
Publicado: (2023)
Real-CATS: A Practical Training Ground for Emerging Research on Cryptocurrency Cybercrime Detection
por: Shi, Jiadong, et al.
Publicado: (2025)
por: Shi, Jiadong, et al.
Publicado: (2025)
Rigidity of proper holomorphic maps between nonequidimensional Fock–Bargmann–Hartogs domains
por: Guicong Su, et al.
Publicado: (2025)
por: Guicong Su, et al.
Publicado: (2025)
Investigation of theoretical infrared spectra on microwave dielectric properties of Re 2 O 3 (Re = Ho, Er, Tm) ceramics
por: Shu‐Yang Ma, et al.
Publicado: (2024)
por: Shu‐Yang Ma, et al.
Publicado: (2024)
Integrating LSTM and BERT for Long-Sequence Data Analysis in Intelligent Tutoring Systems
por: Li, Zhaoxing, et al.
Publicado: (2024)
por: Li, Zhaoxing, et al.
Publicado: (2024)
InterveneBench: Benchmarking LLMs for Intervention Reasoning and Causal Study Design in Real Social Systems
por: Shi, Shaojie, et al.
Publicado: (2026)
por: Shi, Shaojie, et al.
Publicado: (2026)
Harnessing Large Language Models for Disaster Management: A Survey
por: Lei, Zhenyu, et al.
Publicado: (2025)
por: Lei, Zhenyu, et al.
Publicado: (2025)
Switching exploration modes in human mobility
por: Zhong, Lu, et al.
Publicado: (2025)
por: Zhong, Lu, et al.
Publicado: (2025)
Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench
por: Wang, Tianyu, et al.
Publicado: (2026)
por: Wang, Tianyu, et al.
Publicado: (2026)
Aluminum nitride‐based ceramics with excellent thermal shock resistances
por: Zhongyan Wang, et al.
Publicado: (2024)
por: Zhongyan Wang, et al.
Publicado: (2024)
Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent Approach
por: Wu, Yuchen, et al.
Publicado: (2025)
por: Wu, Yuchen, et al.
Publicado: (2025)
Hierarchical Reinforcement Learning for Cooperative Air-Ground Delivery in Urban System
por: Lei, Songxin, et al.
Publicado: (2026)
por: Lei, Songxin, et al.
Publicado: (2026)
Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation
por: Shen, Hanwen, et al.
Publicado: (2026)
por: Shen, Hanwen, et al.
Publicado: (2026)
Information Cocoons on Social Media: Why and How Should the Government Regulate Algorithms
por: Yang, Wen
Publicado: (2024)
por: Yang, Wen
Publicado: (2024)
AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios
por: Mou, Xinyi, et al.
Publicado: (2024)
por: Mou, Xinyi, et al.
Publicado: (2024)
Ejemplares similares
-
YuLan-OneSim: Towards the Next Generation of Social Simulator with Large Language Models
por: Wang, Lei, et al.
Publicado: (2025) -
How Far Are LLMs from Believable AI? A Benchmark for Evaluating the Believability of Human Behavior Simulation
por: Xiao, Yang, et al.
Publicado: (2023) -
LLM Agents as Social Scientists: A Human-AI Collaborative Platform for Social Science Automation
por: Wang, Lei, et al.
Publicado: (2026) -
When Assurance Undermines Intelligence: The Efficiency Costs of Data Governance in AI-Enabled Labor Markets
por: Chen, Lei, et al.
Publicado: (2025) -
Benchmarking LLMs for Political Science: A United Nations Perspective
por: Liang, Yueqing, et al.
Publicado: (2025)