From Model-centered to Human-Centered: Revision Distance as a Metric for Text Evaluation in LLMs-based Applications

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ma, Yongqiang, Qing, Lizhi, Liu, Jiawei, Kang, Yangyang, Zhang, Yue, Lu, Wei, Liu, Xiaozhong, Cheng, Qikai
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913309876289536
author Ma, Yongqiang
Qing, Lizhi
Liu, Jiawei
Kang, Yangyang
Zhang, Yue
Lu, Wei
Liu, Xiaozhong
Cheng, Qikai
author_facet Ma, Yongqiang
Qing, Lizhi
Liu, Jiawei
Kang, Yangyang
Zhang, Yue
Lu, Wei
Liu, Xiaozhong
Cheng, Qikai
contents Evaluating large language models (LLMs) is fundamental, particularly in the context of practical applications. Conventional evaluation methods, typically designed primarily for LLM development, yield numerical scores that ignore the user experience. Therefore, our study shifts the focus from model-centered to human-centered evaluation in the context of AI-powered writing assistance applications. Our proposed metric, termed ``Revision Distance,'' utilizes LLMs to suggest revision edits that mimic the human writing process. It is determined by counting the revision edits generated by LLMs. Benefiting from the generated revision edit details, our metric can provide a self-explained text evaluation result in a human-understandable manner beyond the context-independent score. Our results show that for the easy-writing task, ``Revision Distance'' is consistent with established metrics (ROUGE, Bert-score, and GPT-score), but offers more insightful, detailed feedback and better distinguishes between texts. Moreover, in the context of challenging academic writing tasks, our metric still delivers reliable evaluations where other metrics tend to struggle. Furthermore, our metric also holds significant potential for scenarios lacking reference texts.
format Preprint
id arxiv_https___arxiv_org_abs_2404_07108
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle From Model-centered to Human-Centered: Revision Distance as a Metric for Text Evaluation in LLMs-based Applications
Ma, Yongqiang
Qing, Lizhi
Liu, Jiawei
Kang, Yangyang
Zhang, Yue
Lu, Wei
Liu, Xiaozhong
Cheng, Qikai
Computation and Language
Information Retrieval
Evaluating large language models (LLMs) is fundamental, particularly in the context of practical applications. Conventional evaluation methods, typically designed primarily for LLM development, yield numerical scores that ignore the user experience. Therefore, our study shifts the focus from model-centered to human-centered evaluation in the context of AI-powered writing assistance applications. Our proposed metric, termed ``Revision Distance,'' utilizes LLMs to suggest revision edits that mimic the human writing process. It is determined by counting the revision edits generated by LLMs. Benefiting from the generated revision edit details, our metric can provide a self-explained text evaluation result in a human-understandable manner beyond the context-independent score. Our results show that for the easy-writing task, ``Revision Distance'' is consistent with established metrics (ROUGE, Bert-score, and GPT-score), but offers more insightful, detailed feedback and better distinguishes between texts. Moreover, in the context of challenging academic writing tasks, our metric still delivers reliable evaluations where other metrics tend to struggle. Furthermore, our metric also holds significant potential for scenarios lacking reference texts.
title From Model-centered to Human-Centered: Revision Distance as a Metric for Text Evaluation in LLMs-based Applications
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2404.07108