On Synthesizing Data for Context Attribution in Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Radevski, Gorjan, Gashteovski, Kiril, Syed, Shahbaz, Malon, Christopher, Nicolas, Sebastien, Hung, Chia-Chien, Sztyler, Timo, Heußer, Verena, Rim, Wiem Ben, Enomoto, Masafumi, Takeoka, Kunihiro, Oyamada, Masafumi, Glavaš, Goran, Lawrence, Carolin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912431263973376
author Radevski, Gorjan
Gashteovski, Kiril
Syed, Shahbaz
Malon, Christopher
Nicolas, Sebastien
Hung, Chia-Chien
Sztyler, Timo
Heußer, Verena
Rim, Wiem Ben
Enomoto, Masafumi
Takeoka, Kunihiro
Oyamada, Masafumi
Glavaš, Goran
Lawrence, Carolin
author_facet Radevski, Gorjan
Gashteovski, Kiril
Syed, Shahbaz
Malon, Christopher
Nicolas, Sebastien
Hung, Chia-Chien
Sztyler, Timo
Heußer, Verena
Rim, Wiem Ben
Enomoto, Masafumi
Takeoka, Kunihiro
Oyamada, Masafumi
Glavaš, Goran
Lawrence, Carolin
contents Question Answering (QA) accounts for a significant portion of LLM usage "in the wild". However, LLMs sometimes produce false or misleading responses, also known as "hallucinations". Therefore, grounding the generated answers in contextually provided information -- i.e., providing evidence for the generated text -- is paramount for LLMs' trustworthiness. Providing this information is the task of context attribution. In this paper, we systematically study LLM-based approaches for this task, namely we investigate (i) zero-shot inference, (ii) LLM ensembling, and (iii) fine-tuning of small LMs on synthetic data generated by larger LLMs. Our key contribution is SynQA: a novel generative strategy for synthesizing context attribution data. Given selected context sentences, an LLM generates QA pairs that are supported by these sentences. This leverages LLMs' natural strengths in text generation while ensuring clear attribution paths in the synthetic training data. We show that the attribution data synthesized via SynQA is highly effective for fine-tuning small LMs for context attribution in different QA tasks and domains. Finally, with a user study, we validate the usefulness of small LMs (fine-tuned on synthetic data from SynQA) in context attribution for QA.
format Preprint
id arxiv_https___arxiv_org_abs_2504_05317
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On Synthesizing Data for Context Attribution in Question Answering
Radevski, Gorjan
Gashteovski, Kiril
Syed, Shahbaz
Malon, Christopher
Nicolas, Sebastien
Hung, Chia-Chien
Sztyler, Timo
Heußer, Verena
Rim, Wiem Ben
Enomoto, Masafumi
Takeoka, Kunihiro
Oyamada, Masafumi
Glavaš, Goran
Lawrence, Carolin
Information Retrieval
Artificial Intelligence
Computation and Language
Machine Learning
Question Answering (QA) accounts for a significant portion of LLM usage "in the wild". However, LLMs sometimes produce false or misleading responses, also known as "hallucinations". Therefore, grounding the generated answers in contextually provided information -- i.e., providing evidence for the generated text -- is paramount for LLMs' trustworthiness. Providing this information is the task of context attribution. In this paper, we systematically study LLM-based approaches for this task, namely we investigate (i) zero-shot inference, (ii) LLM ensembling, and (iii) fine-tuning of small LMs on synthetic data generated by larger LLMs. Our key contribution is SynQA: a novel generative strategy for synthesizing context attribution data. Given selected context sentences, an LLM generates QA pairs that are supported by these sentences. This leverages LLMs' natural strengths in text generation while ensuring clear attribution paths in the synthetic training data. We show that the attribution data synthesized via SynQA is highly effective for fine-tuning small LMs for context attribution in different QA tasks and domains. Finally, with a user study, we validate the usefulness of small LMs (fine-tuned on synthetic data from SynQA) in context attribution for QA.
title On Synthesizing Data for Context Attribution in Question Answering
topic Information Retrieval
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2504.05317