Probing LLM World Models: Enhancing Guesstimation with Wisdom of Crowds Decoding
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866916970449862656 |
|---|---|
| author | Chuang, Yun-Shiuan Narendran, Sameer Harlalka, Nikunj Cheung, Alexander Gao, Sizhe Suresh, Siddharth Hu, Junjie Rogers, Timothy T. |
| author_facet | Chuang, Yun-Shiuan Narendran, Sameer Harlalka, Nikunj Cheung, Alexander Gao, Sizhe Suresh, Siddharth Hu, Junjie Rogers, Timothy T. |
| contents | Guesstimation -- the task of making approximate quantitative estimates about objects or events -- is a common real-world skill, yet remains underexplored in large language model (LLM) research. We introduce three guesstimation datasets: MARBLES, FUTURE, and ELECPRED, spanning physical estimation (e.g., how many marbles fit in a cup) to abstract predictions (e.g., the 2024 U.S. presidential election). Inspired by the social science concept of Wisdom of Crowds (WOC)- where the median of multiple estimates improves accuracy-we propose WOC decoding for LLMs. We replicate WOC effects in human participants and find that LLMs exhibit similar benefits: median aggregation across sampled responses consistently improves accuracy over greedy decoding, self-consistency decoding, and mean decoding. This suggests that LLMs encode a world model that supports approximate reasoning. Our results position guesstimation as a useful probe of LLM world knowledge and highlight WOC decoding as a strategy for enhancing LLM guesstimation performance on real-world tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_17310 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Probing LLM World Models: Enhancing Guesstimation with Wisdom of Crowds Decoding Chuang, Yun-Shiuan Narendran, Sameer Harlalka, Nikunj Cheung, Alexander Gao, Sizhe Suresh, Siddharth Hu, Junjie Rogers, Timothy T. Artificial Intelligence Human-Computer Interaction Guesstimation -- the task of making approximate quantitative estimates about objects or events -- is a common real-world skill, yet remains underexplored in large language model (LLM) research. We introduce three guesstimation datasets: MARBLES, FUTURE, and ELECPRED, spanning physical estimation (e.g., how many marbles fit in a cup) to abstract predictions (e.g., the 2024 U.S. presidential election). Inspired by the social science concept of Wisdom of Crowds (WOC)- where the median of multiple estimates improves accuracy-we propose WOC decoding for LLMs. We replicate WOC effects in human participants and find that LLMs exhibit similar benefits: median aggregation across sampled responses consistently improves accuracy over greedy decoding, self-consistency decoding, and mean decoding. This suggests that LLMs encode a world model that supports approximate reasoning. Our results position guesstimation as a useful probe of LLM world knowledge and highlight WOC decoding as a strategy for enhancing LLM guesstimation performance on real-world tasks. |
| title | Probing LLM World Models: Enhancing Guesstimation with Wisdom of Crowds Decoding |
| topic | Artificial Intelligence Human-Computer Interaction |
| url | https://arxiv.org/abs/2501.17310 |