Probing LLM World Models: Enhancing Guesstimation with Wisdom of Crowds Decoding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chuang, Yun-Shiuan, Narendran, Sameer, Harlalka, Nikunj, Cheung, Alexander, Gao, Sizhe, Suresh, Siddharth, Hu, Junjie, Rogers, Timothy T.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916970449862656
author Chuang, Yun-Shiuan
Narendran, Sameer
Harlalka, Nikunj
Cheung, Alexander
Gao, Sizhe
Suresh, Siddharth
Hu, Junjie
Rogers, Timothy T.
author_facet Chuang, Yun-Shiuan
Narendran, Sameer
Harlalka, Nikunj
Cheung, Alexander
Gao, Sizhe
Suresh, Siddharth
Hu, Junjie
Rogers, Timothy T.
contents Guesstimation -- the task of making approximate quantitative estimates about objects or events -- is a common real-world skill, yet remains underexplored in large language model (LLM) research. We introduce three guesstimation datasets: MARBLES, FUTURE, and ELECPRED, spanning physical estimation (e.g., how many marbles fit in a cup) to abstract predictions (e.g., the 2024 U.S. presidential election). Inspired by the social science concept of Wisdom of Crowds (WOC)- where the median of multiple estimates improves accuracy-we propose WOC decoding for LLMs. We replicate WOC effects in human participants and find that LLMs exhibit similar benefits: median aggregation across sampled responses consistently improves accuracy over greedy decoding, self-consistency decoding, and mean decoding. This suggests that LLMs encode a world model that supports approximate reasoning. Our results position guesstimation as a useful probe of LLM world knowledge and highlight WOC decoding as a strategy for enhancing LLM guesstimation performance on real-world tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2501_17310
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Probing LLM World Models: Enhancing Guesstimation with Wisdom of Crowds Decoding
Chuang, Yun-Shiuan
Narendran, Sameer
Harlalka, Nikunj
Cheung, Alexander
Gao, Sizhe
Suresh, Siddharth
Hu, Junjie
Rogers, Timothy T.
Artificial Intelligence
Human-Computer Interaction
Guesstimation -- the task of making approximate quantitative estimates about objects or events -- is a common real-world skill, yet remains underexplored in large language model (LLM) research. We introduce three guesstimation datasets: MARBLES, FUTURE, and ELECPRED, spanning physical estimation (e.g., how many marbles fit in a cup) to abstract predictions (e.g., the 2024 U.S. presidential election). Inspired by the social science concept of Wisdom of Crowds (WOC)- where the median of multiple estimates improves accuracy-we propose WOC decoding for LLMs. We replicate WOC effects in human participants and find that LLMs exhibit similar benefits: median aggregation across sampled responses consistently improves accuracy over greedy decoding, self-consistency decoding, and mean decoding. This suggests that LLMs encode a world model that supports approximate reasoning. Our results position guesstimation as a useful probe of LLM world knowledge and highlight WOC decoding as a strategy for enhancing LLM guesstimation performance on real-world tasks.
title Probing LLM World Models: Enhancing Guesstimation with Wisdom of Crowds Decoding
topic Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2501.17310