Listen to the Context: Towards Faithful Large Language Models for Retrieval Augmented Generation on Climate Questions

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Thulke, David, Kemmler, Jakob, Dugast, Christian, Ney, Hermann
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909618800689152
author Thulke, David
Kemmler, Jakob
Dugast, Christian
Ney, Hermann
author_facet Thulke, David
Kemmler, Jakob
Dugast, Christian
Ney, Hermann
contents Large language models that use retrieval augmented generation have the potential to unlock valuable knowledge for researchers, policymakers, and the public by making long and technical climate-related documents more accessible. While this approach can help alleviate factual hallucinations by relying on retrieved passages as additional context, its effectiveness depends on whether the model's output remains faithful to these passages. To address this, we explore the automatic assessment of faithfulness of different models in this setting. We then focus on ClimateGPT, a large language model specialised in climate science, to examine which factors in its instruction fine-tuning impact the model's faithfulness. By excluding unfaithful subsets of the model's training data, we develop ClimateGPT Faithful+, which achieves an improvement in faithfulness from 30% to 57% in supported atomic claims according to our automatic metric.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15633
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Listen to the Context: Towards Faithful Large Language Models for Retrieval Augmented Generation on Climate Questions
Thulke, David
Kemmler, Jakob
Dugast, Christian
Ney, Hermann
Computation and Language
Artificial Intelligence
Machine Learning
Large language models that use retrieval augmented generation have the potential to unlock valuable knowledge for researchers, policymakers, and the public by making long and technical climate-related documents more accessible. While this approach can help alleviate factual hallucinations by relying on retrieved passages as additional context, its effectiveness depends on whether the model's output remains faithful to these passages. To address this, we explore the automatic assessment of faithfulness of different models in this setting. We then focus on ClimateGPT, a large language model specialised in climate science, to examine which factors in its instruction fine-tuning impact the model's faithfulness. By excluding unfaithful subsets of the model's training data, we develop ClimateGPT Faithful+, which achieves an improvement in faithfulness from 30% to 57% in supported atomic claims according to our automatic metric.
title Listen to the Context: Towards Faithful Large Language Models for Retrieval Augmented Generation on Climate Questions
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.15633