Medal Matters: Probing LLMs' Failure Cases Through Olympic Rankings

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Choi, Juhwan, Yu, Seunguk, Yun, JungMin, Kim, YoungBin
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908780117098496
author Choi, Juhwan
Yu, Seunguk
Yun, JungMin
Kim, YoungBin
author_facet Choi, Juhwan
Yu, Seunguk
Yun, JungMin
Kim, YoungBin
contents Large language models (LLMs) have achieved remarkable success in natural language processing tasks, yet their internal knowledge structures remain poorly understood. This study examines these structures through the lens of historical Olympic medal tallies, evaluating LLMs on two tasks: (1) retrieving medal counts for specific teams and (2) identifying rankings of each team. While state-of-the-art LLMs excel in recalling medal counts, they struggle with providing rankings, highlighting a key difference between their knowledge organization and human reasoning. These findings shed light on the limitations of LLMs' internal knowledge integration and suggest directions for improvement. To facilitate further research, we release our code, dataset, and model outputs.
format Preprint
id arxiv_https___arxiv_org_abs_2409_06518
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Medal Matters: Probing LLMs' Failure Cases Through Olympic Rankings
Choi, Juhwan
Yu, Seunguk
Yun, JungMin
Kim, YoungBin
Computation and Language
Artificial Intelligence
Large language models (LLMs) have achieved remarkable success in natural language processing tasks, yet their internal knowledge structures remain poorly understood. This study examines these structures through the lens of historical Olympic medal tallies, evaluating LLMs on two tasks: (1) retrieving medal counts for specific teams and (2) identifying rankings of each team. While state-of-the-art LLMs excel in recalling medal counts, they struggle with providing rankings, highlighting a key difference between their knowledge organization and human reasoning. These findings shed light on the limitations of LLMs' internal knowledge integration and suggest directions for improvement. To facilitate further research, we release our code, dataset, and model outputs.
title Medal Matters: Probing LLMs' Failure Cases Through Olympic Rankings
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2409.06518