ConSens: Assessing context grounding in open-book question answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vankov, Ivan, Ivanov, Matyo, Correia, Adriana, Botev, Victor
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916715671060480
author Vankov, Ivan
Ivanov, Matyo
Correia, Adriana
Botev, Victor
author_facet Vankov, Ivan
Ivanov, Matyo
Correia, Adriana
Botev, Victor
contents Large Language Models (LLMs) have demonstrated considerable success in open-book question answering (QA), where the task requires generating answers grounded in a provided external context. A critical challenge in open-book QA is to ensure that model responses are based on the provided context rather than its parametric knowledge, which can be outdated, incomplete, or incorrect. Existing evaluation methods, primarily based on the LLM-as-a-judge approach, face significant limitations, including biases, scalability issues, and dependence on costly external systems. To address these challenges, we propose a novel metric that contrasts the perplexity of the model response under two conditions: when the context is provided and when it is not. The resulting score quantifies the extent to which the model's answer relies on the provided context. The validity of this metric is demonstrated through a series of experiments that show its effectiveness in identifying whether a given answer is grounded in the provided context. Unlike existing approaches, this metric is computationally efficient, interpretable, and adaptable to various use cases, offering a scalable and practical solution to assess context utilization in open-book QA systems.
format Preprint
id arxiv_https___arxiv_org_abs_2505_00065
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ConSens: Assessing context grounding in open-book question answering
Vankov, Ivan
Ivanov, Matyo
Correia, Adriana
Botev, Victor
Computation and Language
Machine Learning
Large Language Models (LLMs) have demonstrated considerable success in open-book question answering (QA), where the task requires generating answers grounded in a provided external context. A critical challenge in open-book QA is to ensure that model responses are based on the provided context rather than its parametric knowledge, which can be outdated, incomplete, or incorrect. Existing evaluation methods, primarily based on the LLM-as-a-judge approach, face significant limitations, including biases, scalability issues, and dependence on costly external systems. To address these challenges, we propose a novel metric that contrasts the perplexity of the model response under two conditions: when the context is provided and when it is not. The resulting score quantifies the extent to which the model's answer relies on the provided context. The validity of this metric is demonstrated through a series of experiments that show its effectiveness in identifying whether a given answer is grounded in the provided context. Unlike existing approaches, this metric is computationally efficient, interpretable, and adaptable to various use cases, offering a scalable and practical solution to assess context utilization in open-book QA systems.
title ConSens: Assessing context grounding in open-book question answering
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.00065