QUILL: Quotation Generation Enhancement of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Jin, Zhang, Bowei, He, Qianyu, Liang, Jiaqing, Wei, Feng, Chen, Jinglei, Liang, Zujie, Yang, Deqing, Xiao, Yanghua
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915162342031360
author Xiao, Jin
Zhang, Bowei
He, Qianyu
Liang, Jiaqing
Wei, Feng
Chen, Jinglei
Liang, Zujie
Yang, Deqing
Xiao, Yanghua
author_facet Xiao, Jin
Zhang, Bowei
He, Qianyu
Liang, Jiaqing
Wei, Feng
Chen, Jinglei
Liang, Zujie
Yang, Deqing
Xiao, Yanghua
contents While Large language models (LLMs) have become excellent writing assistants, they still struggle with quotation generation. This is because they either hallucinate when providing factual quotations or fail to provide quotes that exceed human expectations. To bridge the gap, we systematically study how to evaluate and improve LLMs' performance in quotation generation tasks. We first establish a holistic and automatic evaluation system for quotation generation task, which consists of five criteria each with corresponding automatic metric. To improve the LLMs' quotation generation abilities, we construct a bilingual knowledge base that is broad in scope and rich in dimensions, containing up to 32,022 quotes. Moreover, guided by our critiria, we further design a quotation-specific metric to rerank the retrieved quotations from the knowledge base. Extensive experiments show that our metrics strongly correlate with human preferences. Existing LLMs struggle to generate desired quotes, but our quotation knowledge base and reranking metric help narrow this gap. Our dataset and code are publicly available at https://github.com/GraceXiaoo/QUILL.
format Preprint
id arxiv_https___arxiv_org_abs_2411_03675
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle QUILL: Quotation Generation Enhancement of Large Language Models
Xiao, Jin
Zhang, Bowei
He, Qianyu
Liang, Jiaqing
Wei, Feng
Chen, Jinglei
Liang, Zujie
Yang, Deqing
Xiao, Yanghua
Computation and Language
Artificial Intelligence
While Large language models (LLMs) have become excellent writing assistants, they still struggle with quotation generation. This is because they either hallucinate when providing factual quotations or fail to provide quotes that exceed human expectations. To bridge the gap, we systematically study how to evaluate and improve LLMs' performance in quotation generation tasks. We first establish a holistic and automatic evaluation system for quotation generation task, which consists of five criteria each with corresponding automatic metric. To improve the LLMs' quotation generation abilities, we construct a bilingual knowledge base that is broad in scope and rich in dimensions, containing up to 32,022 quotes. Moreover, guided by our critiria, we further design a quotation-specific metric to rerank the retrieved quotations from the knowledge base. Extensive experiments show that our metrics strongly correlate with human preferences. Existing LLMs struggle to generate desired quotes, but our quotation knowledge base and reranking metric help narrow this gap. Our dataset and code are publicly available at https://github.com/GraceXiaoo/QUILL.
title QUILL: Quotation Generation Enhancement of Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2411.03675