Retrieval-Constrained Decoding Reveals Underestimated Parametric Knowledge in Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hamdani, Rajaa El, Haffoudhi, Samy, Holzenberger, Nils, Suchanek, Fabian, Bonald, Thomas, Malliaros, Fragkiskos D.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914061206159360
author Hamdani, Rajaa El
Haffoudhi, Samy
Holzenberger, Nils
Suchanek, Fabian
Bonald, Thomas
Malliaros, Fragkiskos D.
author_facet Hamdani, Rajaa El
Haffoudhi, Samy
Holzenberger, Nils
Suchanek, Fabian
Bonald, Thomas
Malliaros, Fragkiskos D.
contents Language models (LMs) encode substantial factual knowledge, but often produce answers judged as incorrect. We hypothesize that many of these answers are actually correct, but are expressed in alternative surface forms that are dismissed due to an overly strict evaluation, leading to an underestimation of models' parametric knowledge. We propose Retrieval-Constrained Decoding (RCD), a decoding strategy that restricts model outputs to unique surface forms. We introduce YAGO-QA, a dataset of 19,137 general knowledge questions. Evaluating open-source LMs from 135M to 70B parameters, we show that standard decoding undervalues their knowledge. For instance, Llama-3.1-70B scores only 32.3% F1 with vanilla decoding but 46.0% with RCD. Similarly, Llama-3.1-8B reaches 33.0% with RCD, outperforming the larger model under vanilla decoding. We publicly share the code and dataset at https://github.com/Rajjaa/disambiguated-LLM.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23417
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Retrieval-Constrained Decoding Reveals Underestimated Parametric Knowledge in Language Models
Hamdani, Rajaa El
Haffoudhi, Samy
Holzenberger, Nils
Suchanek, Fabian
Bonald, Thomas
Malliaros, Fragkiskos D.
Computation and Language
Artificial Intelligence
Language models (LMs) encode substantial factual knowledge, but often produce answers judged as incorrect. We hypothesize that many of these answers are actually correct, but are expressed in alternative surface forms that are dismissed due to an overly strict evaluation, leading to an underestimation of models' parametric knowledge. We propose Retrieval-Constrained Decoding (RCD), a decoding strategy that restricts model outputs to unique surface forms. We introduce YAGO-QA, a dataset of 19,137 general knowledge questions. Evaluating open-source LMs from 135M to 70B parameters, we show that standard decoding undervalues their knowledge. For instance, Llama-3.1-70B scores only 32.3% F1 with vanilla decoding but 46.0% with RCD. Similarly, Llama-3.1-8B reaches 33.0% with RCD, outperforming the larger model under vanilla decoding. We publicly share the code and dataset at https://github.com/Rajjaa/disambiguated-LLM.
title Retrieval-Constrained Decoding Reveals Underestimated Parametric Knowledge in Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.23417