CrowdLLM: Building LLM-Based Digital Populations Augmented with Generative Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lin, Ryan Feng, Tian, Keyu, Zheng, Hanming, Zhang, Congjing, Zeng, Li, Huang, Shuai
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914254320304128
author Lin, Ryan Feng
Tian, Keyu
Zheng, Hanming
Zhang, Congjing
Zeng, Li
Huang, Shuai
author_facet Lin, Ryan Feng
Tian, Keyu
Zheng, Hanming
Zhang, Congjing
Zeng, Li
Huang, Shuai
contents The emergence of large language models (LLMs) has sparked much interest in creating LLM-based digital populations that can be applied to many applications such as social simulation, crowdsourcing, marketing, and recommendation systems. A digital population can reduce the cost of recruiting human participants and alleviate many concerns related to human subject study. However, research has found that most of the existing works rely solely on LLMs and could not sufficiently capture the accuracy and diversity of a real human population. To address this limitation, we propose CrowdLLM that integrates pretrained LLMs and generative models to enhance the diversity and fidelity of the digital population. We conduct theoretical analysis of CrowdLLM regarding its great potential in creating cost-effective, sufficiently representative, scalable digital populations that can match the quality of a real crowd. Comprehensive experiments are also conducted across multiple domains (e.g., crowdsourcing, voting, user rating) and simulation studies which demonstrate that CrowdLLM achieves promising performance in both accuracy and distributional fidelity to human data.
format Preprint
id arxiv_https___arxiv_org_abs_2512_07890
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CrowdLLM: Building LLM-Based Digital Populations Augmented with Generative Models
Lin, Ryan Feng
Tian, Keyu
Zheng, Hanming
Zhang, Congjing
Zeng, Li
Huang, Shuai
Multiagent Systems
Artificial Intelligence
Machine Learning
Methodology
The emergence of large language models (LLMs) has sparked much interest in creating LLM-based digital populations that can be applied to many applications such as social simulation, crowdsourcing, marketing, and recommendation systems. A digital population can reduce the cost of recruiting human participants and alleviate many concerns related to human subject study. However, research has found that most of the existing works rely solely on LLMs and could not sufficiently capture the accuracy and diversity of a real human population. To address this limitation, we propose CrowdLLM that integrates pretrained LLMs and generative models to enhance the diversity and fidelity of the digital population. We conduct theoretical analysis of CrowdLLM regarding its great potential in creating cost-effective, sufficiently representative, scalable digital populations that can match the quality of a real crowd. Comprehensive experiments are also conducted across multiple domains (e.g., crowdsourcing, voting, user rating) and simulation studies which demonstrate that CrowdLLM achieves promising performance in both accuracy and distributional fidelity to human data.
title CrowdLLM: Building LLM-Based Digital Populations Augmented with Generative Models
topic Multiagent Systems
Artificial Intelligence
Machine Learning
Methodology
url https://arxiv.org/abs/2512.07890