Assessing LLM Reasoning Steps via Principal Knowledge Grounding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hwang, Hyeon, Cho, Yewon, Yoon, Chanwoong, Park, Yein, Song, Minju, Lee, Kyungjae, Kim, Gangwoo, Kang, Jaewoo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911246526185472
author Hwang, Hyeon
Cho, Yewon
Yoon, Chanwoong
Park, Yein
Song, Minju
Lee, Kyungjae
Kim, Gangwoo
Kang, Jaewoo
author_facet Hwang, Hyeon
Cho, Yewon
Yoon, Chanwoong
Park, Yein
Song, Minju
Lee, Kyungjae
Kim, Gangwoo
Kang, Jaewoo
contents Step-by-step reasoning has become a standard approach for large language models (LLMs) to tackle complex tasks. While this paradigm has proven effective, it raises a fundamental question: How can we verify that an LLM's reasoning is accurately grounded in knowledge? To address this question, we introduce a novel evaluation suite that systematically assesses the knowledge grounding of intermediate reasoning. Our framework comprises three key components. (1) Principal Knowledge Collection, a large-scale repository of atomic knowledge essential for reasoning. Based on the collection, we propose (2) knowledge-grounded evaluation metrics designed to measure how well models recall and apply prerequisite knowledge in reasoning. These metrics are computed by our (3) evaluator LLM, a lightweight model optimized for cost-effective and reliable metric computation. Our evaluation suite demonstrates remarkable effectiveness in identifying missing or misapplied knowledge elements, providing crucial insights for uncovering fundamental reasoning deficiencies in LLMs. Beyond evaluation, we demonstrate how these metrics can be integrated into preference optimization, showcasing further applications of knowledge-grounded evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00879
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Assessing LLM Reasoning Steps via Principal Knowledge Grounding
Hwang, Hyeon
Cho, Yewon
Yoon, Chanwoong
Park, Yein
Song, Minju
Lee, Kyungjae
Kim, Gangwoo
Kang, Jaewoo
Computation and Language
Artificial Intelligence
Machine Learning
Step-by-step reasoning has become a standard approach for large language models (LLMs) to tackle complex tasks. While this paradigm has proven effective, it raises a fundamental question: How can we verify that an LLM's reasoning is accurately grounded in knowledge? To address this question, we introduce a novel evaluation suite that systematically assesses the knowledge grounding of intermediate reasoning. Our framework comprises three key components. (1) Principal Knowledge Collection, a large-scale repository of atomic knowledge essential for reasoning. Based on the collection, we propose (2) knowledge-grounded evaluation metrics designed to measure how well models recall and apply prerequisite knowledge in reasoning. These metrics are computed by our (3) evaluator LLM, a lightweight model optimized for cost-effective and reliable metric computation. Our evaluation suite demonstrates remarkable effectiveness in identifying missing or misapplied knowledge elements, providing crucial insights for uncovering fundamental reasoning deficiencies in LLMs. Beyond evaluation, we demonstrate how these metrics can be integrated into preference optimization, showcasing further applications of knowledge-grounded evaluation.
title Assessing LLM Reasoning Steps via Principal Knowledge Grounding
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2511.00879