Medal Matters: Probing LLMs' Failure Cases Through Olympic Rankings
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866908780117098496 |
|---|---|
| author | Choi, Juhwan Yu, Seunguk Yun, JungMin Kim, YoungBin |
| author_facet | Choi, Juhwan Yu, Seunguk Yun, JungMin Kim, YoungBin |
| contents | Large language models (LLMs) have achieved remarkable success in natural language processing tasks, yet their internal knowledge structures remain poorly understood. This study examines these structures through the lens of historical Olympic medal tallies, evaluating LLMs on two tasks: (1) retrieving medal counts for specific teams and (2) identifying rankings of each team. While state-of-the-art LLMs excel in recalling medal counts, they struggle with providing rankings, highlighting a key difference between their knowledge organization and human reasoning. These findings shed light on the limitations of LLMs' internal knowledge integration and suggest directions for improvement. To facilitate further research, we release our code, dataset, and model outputs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_06518 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Medal Matters: Probing LLMs' Failure Cases Through Olympic Rankings Choi, Juhwan Yu, Seunguk Yun, JungMin Kim, YoungBin Computation and Language Artificial Intelligence Large language models (LLMs) have achieved remarkable success in natural language processing tasks, yet their internal knowledge structures remain poorly understood. This study examines these structures through the lens of historical Olympic medal tallies, evaluating LLMs on two tasks: (1) retrieving medal counts for specific teams and (2) identifying rankings of each team. While state-of-the-art LLMs excel in recalling medal counts, they struggle with providing rankings, highlighting a key difference between their knowledge organization and human reasoning. These findings shed light on the limitations of LLMs' internal knowledge integration and suggest directions for improvement. To facilitate further research, we release our code, dataset, and model outputs. |
| title | Medal Matters: Probing LLMs' Failure Cases Through Olympic Rankings |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2409.06518 |