Transparentize the Internal and External Knowledge Utilization in LLMs with Trustworthy Citation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Jiajun, Zhou, Tong, Chen, Yubo, Qiu, Delai, Liu, Shengping, Liu, Kang, Zhao, Jun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910915368058880
author Shen, Jiajun
Zhou, Tong
Chen, Yubo
Qiu, Delai
Liu, Shengping
Liu, Kang
Zhao, Jun
author_facet Shen, Jiajun
Zhou, Tong
Chen, Yubo
Qiu, Delai
Liu, Shengping
Liu, Kang
Zhao, Jun
contents While hallucinations of large language models could been alleviated through retrieval-augmented generation and citation generation, how the model utilizes internal knowledge is still opaque, and the trustworthiness of its generated answers remains questionable. In this work, we introduce Context-Prior Augmented Citation Generation task, requiring models to generate citations considering both external and internal knowledge while providing trustworthy references, with 5 evaluation metrics focusing on 3 aspects: answer helpfulness, citation faithfulness, and trustworthiness. We introduce RAEL, the paradigm for our task, and also design INTRALIGN, an integrated method containing customary data generation and an alignment algorithm. Our experimental results show that our method achieves a better cross-scenario performance with regard to other baselines. Our extended experiments further reveal that retrieval quality, question types, and model knowledge have considerable influence on the trustworthiness in citation generation.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14856
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Transparentize the Internal and External Knowledge Utilization in LLMs with Trustworthy Citation
Shen, Jiajun
Zhou, Tong
Chen, Yubo
Qiu, Delai
Liu, Shengping
Liu, Kang
Zhao, Jun
Computation and Language
While hallucinations of large language models could been alleviated through retrieval-augmented generation and citation generation, how the model utilizes internal knowledge is still opaque, and the trustworthiness of its generated answers remains questionable. In this work, we introduce Context-Prior Augmented Citation Generation task, requiring models to generate citations considering both external and internal knowledge while providing trustworthy references, with 5 evaluation metrics focusing on 3 aspects: answer helpfulness, citation faithfulness, and trustworthiness. We introduce RAEL, the paradigm for our task, and also design INTRALIGN, an integrated method containing customary data generation and an alignment algorithm. Our experimental results show that our method achieves a better cross-scenario performance with regard to other baselines. Our extended experiments further reveal that retrieval quality, question types, and model knowledge have considerable influence on the trustworthiness in citation generation.
title Transparentize the Internal and External Knowledge Utilization in LLMs with Trustworthy Citation
topic Computation and Language
url https://arxiv.org/abs/2504.14856