ConCISE: A Reference-Free Conciseness Evaluation Metric for LLM-Generated Answers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866914387317489664 |
|---|---|
| author | Ghafari, Seyed Mohssen Kol, Ronny Quiroz, Juan C. Luan, Nella Patial, Monika Rupasinghe, Chanaka Wandabwa, Herman Pizzato, Luiz |
| author_facet | Ghafari, Seyed Mohssen Kol, Ronny Quiroz, Juan C. Luan, Nella Patial, Monika Rupasinghe, Chanaka Wandabwa, Herman Pizzato, Luiz |
| contents | Large language models (LLMs) frequently generate responses that are lengthy and verbose, filled with redundant or unnecessary details. This diminishes clarity and user satisfaction, and it increases costs for model developers, especially with well-known proprietary models that charge based on the number of output tokens. In this paper, we introduce a novel reference-free metric for evaluating the conciseness of responses generated by LLMs. Our method quantifies non-essential content without relying on gold standard references and calculates the average of three calculations: i) a compression ratio between the original response and an LLM abstractive summary; ii) a compression ratio between the original response and an LLM extractive summary; and iii) wordremoval compression, where an LLM removes as many non-essential words as possible from the response while preserving its meaning, with the number of tokens removed indicating the conciseness score. Experimental results demonstrate that our proposed metric identifies redundancy in LLM outputs, offering a practical tool for automated evaluation of response brevity in conversational AI systems without the need for ground truth human annotations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_16846 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ConCISE: A Reference-Free Conciseness Evaluation Metric for LLM-Generated Answers Ghafari, Seyed Mohssen Kol, Ronny Quiroz, Juan C. Luan, Nella Patial, Monika Rupasinghe, Chanaka Wandabwa, Herman Pizzato, Luiz Computation and Language Artificial Intelligence I.2.0 Large language models (LLMs) frequently generate responses that are lengthy and verbose, filled with redundant or unnecessary details. This diminishes clarity and user satisfaction, and it increases costs for model developers, especially with well-known proprietary models that charge based on the number of output tokens. In this paper, we introduce a novel reference-free metric for evaluating the conciseness of responses generated by LLMs. Our method quantifies non-essential content without relying on gold standard references and calculates the average of three calculations: i) a compression ratio between the original response and an LLM abstractive summary; ii) a compression ratio between the original response and an LLM extractive summary; and iii) wordremoval compression, where an LLM removes as many non-essential words as possible from the response while preserving its meaning, with the number of tokens removed indicating the conciseness score. Experimental results demonstrate that our proposed metric identifies redundancy in LLM outputs, offering a practical tool for automated evaluation of response brevity in conversational AI systems without the need for ground truth human annotations. |
| title | ConCISE: A Reference-Free Conciseness Evaluation Metric for LLM-Generated Answers |
| topic | Computation and Language Artificial Intelligence I.2.0 |
| url | https://arxiv.org/abs/2511.16846 |