Rate, Explain and Cite (REC): Enhanced Explanation and Attribution in Automatic Evaluation by Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hsu, Aliyah R., Zhu, James, Wang, Zhichao, Bi, Bin, Mehrotra, Shubham, Pentyala, Shiva K., Tan, Katherine, Mao, Xiang-Bo, Omrani, Roshanak, Chaudhuri, Sougata, Radhakrishnan, Regunathan, Asur, Sitaram, Cheng, Claire Na, Yu, Bin
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909616571416576
author Hsu, Aliyah R.
Zhu, James
Wang, Zhichao
Bi, Bin
Mehrotra, Shubham
Pentyala, Shiva K.
Tan, Katherine
Mao, Xiang-Bo
Omrani, Roshanak
Chaudhuri, Sougata
Radhakrishnan, Regunathan
Asur, Sitaram
Cheng, Claire Na
Yu, Bin
author_facet Hsu, Aliyah R.
Zhu, James
Wang, Zhichao
Bi, Bin
Mehrotra, Shubham
Pentyala, Shiva K.
Tan, Katherine
Mao, Xiang-Bo
Omrani, Roshanak
Chaudhuri, Sougata
Radhakrishnan, Regunathan
Asur, Sitaram
Cheng, Claire Na
Yu, Bin
contents LLMs have demonstrated impressive proficiency in generating coherent and high-quality text, making them valuable across a range of text-generation tasks. However, rigorous evaluation of this generated content is crucial, as ensuring its quality remains a significant challenge due to persistent issues such as factual inaccuracies and hallucination. This paper introduces three fine-tuned general-purpose LLM autoevaluators, REC-8B, REC-12B and REC-70B, specifically designed to evaluate generated text across several dimensions: faithfulness, instruction following, coherence, and completeness. These models not only provide ratings for these metrics but also offer detailed explanation and verifiable citation, thereby enhancing trust in the content. Moreover, the models support various citation modes, accommodating different requirements for latency and granularity. Extensive evaluations on diverse benchmarks demonstrate that our general-purpose LLM auto-evaluator, REC-70B, outperforms state-of-the-art LLMs, excelling in content evaluation by delivering better quality explanation and citation with minimal bias. Our REC dataset and models are available at https://github.com/adelaidehsu/REC.
format Preprint
id arxiv_https___arxiv_org_abs_2411_02448
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Rate, Explain and Cite (REC): Enhanced Explanation and Attribution in Automatic Evaluation by Large Language Models
Hsu, Aliyah R.
Zhu, James
Wang, Zhichao
Bi, Bin
Mehrotra, Shubham
Pentyala, Shiva K.
Tan, Katherine
Mao, Xiang-Bo
Omrani, Roshanak
Chaudhuri, Sougata
Radhakrishnan, Regunathan
Asur, Sitaram
Cheng, Claire Na
Yu, Bin
Computation and Language
Artificial Intelligence
LLMs have demonstrated impressive proficiency in generating coherent and high-quality text, making them valuable across a range of text-generation tasks. However, rigorous evaluation of this generated content is crucial, as ensuring its quality remains a significant challenge due to persistent issues such as factual inaccuracies and hallucination. This paper introduces three fine-tuned general-purpose LLM autoevaluators, REC-8B, REC-12B and REC-70B, specifically designed to evaluate generated text across several dimensions: faithfulness, instruction following, coherence, and completeness. These models not only provide ratings for these metrics but also offer detailed explanation and verifiable citation, thereby enhancing trust in the content. Moreover, the models support various citation modes, accommodating different requirements for latency and granularity. Extensive evaluations on diverse benchmarks demonstrate that our general-purpose LLM auto-evaluator, REC-70B, outperforms state-of-the-art LLMs, excelling in content evaluation by delivering better quality explanation and citation with minimal bias. Our REC dataset and models are available at https://github.com/adelaidehsu/REC.
title Rate, Explain and Cite (REC): Enhanced Explanation and Attribution in Automatic Evaluation by Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2411.02448