Reference-free Evaluation Metrics for Text Generation: A Survey
Fuente:
arXiv
Saved in:
| Main Authors: | Ito, Takumi, van Deemter, Kees, Suzuki, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Intrinsic Task-based Evaluation for Referring Expression Generation
by: Chen, Guanyi, et al.
Published: (2024)
by: Chen, Guanyi, et al.
Published: (2024)
The Pitfalls of Defining Hallucination
by: van Deemter, Kees
Published: (2024)
by: van Deemter, Kees
Published: (2024)
My Life in Artificial Intelligence: People, anecdotes, and some lessons learnt
by: van Deemter, Kees
Published: (2025)
by: van Deemter, Kees
Published: (2025)
Computational Modelling of Plurality and Definiteness in Chinese Noun Phrases
by: Liu, Yuqi, et al.
Published: (2024)
by: Liu, Yuqi, et al.
Published: (2024)
Textual Summarisation of Large Sets: Towards a General Approach
by: Kuptavanich, Kittipitch, et al.
Published: (2024)
by: Kuptavanich, Kittipitch, et al.
Published: (2024)
Reliability Crisis of Reference-free Metrics for Grammatical Error Correction
by: Goto, Takumi, et al.
Published: (2025)
by: Goto, Takumi, et al.
Published: (2025)
STEP: Staged Parameter-Efficient Pre-training for Large Language Models
by: Yano, Kazuki, et al.
Published: (2025)
by: Yano, Kazuki, et al.
Published: (2025)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
by: Kasaei, Seyed Amir, et al.
Published: (2025)
by: Kasaei, Seyed Amir, et al.
Published: (2025)
Conceptual Cultural Index: A Metric for Cultural Specificity via Relative Generality
by: Ohashi, Takumi, et al.
Published: (2026)
by: Ohashi, Takumi, et al.
Published: (2026)
Rethinking Evaluation Metrics for Grammatical Error Correction: Why Use a Different Evaluation Process than Human?
by: Goto, Takumi, et al.
Published: (2025)
by: Goto, Takumi, et al.
Published: (2025)
IMPARA-GED: Grammatical Error Detection is Boosting Reference-free Grammatical Error Quality Estimator
by: Sakai, Yusuke, et al.
Published: (2025)
by: Sakai, Yusuke, et al.
Published: (2025)
Automatic Metrics in Natural Language Generation: A Survey of Current Evaluation Practices
by: Schmidtová, Patrícia, et al.
Published: (2024)
by: Schmidtová, Patrícia, et al.
Published: (2024)
Transparent Reference-free Automated Evaluation of Open-Ended User Survey Responses
by: An, Subin, et al.
Published: (2025)
by: An, Subin, et al.
Published: (2025)
Reproducing the Metric-Based Evaluation of a Set of Controllable Text Generation Techniques
by: Lorandi, Michela, et al.
Published: (2024)
by: Lorandi, Michela, et al.
Published: (2024)
Not All Metrics Are Guilty: Improving NLG Evaluation by Diversifying References
by: Tang, Tianyi, et al.
Published: (2023)
by: Tang, Tianyi, et al.
Published: (2023)
Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics?
by: Mukherjee, Sourabrata, et al.
Published: (2025)
by: Mukherjee, Sourabrata, et al.
Published: (2025)
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation
by: Dejl, Adam, et al.
Published: (2025)
by: Dejl, Adam, et al.
Published: (2025)
Evaluating the Utility of Grounding Documents with Reference-Free LLM-based Metrics
by: Hua, Yilun, et al.
Published: (2026)
by: Hua, Yilun, et al.
Published: (2026)
Identifying Reliable Evaluation Metrics for Scientific Text Revision
by: Jourdan, Léane, et al.
Published: (2025)
by: Jourdan, Léane, et al.
Published: (2025)
Hacking Neural Evaluation Metrics with Single Hub Text
by: Deguchi, Hiroyuki, et al.
Published: (2025)
by: Deguchi, Hiroyuki, et al.
Published: (2025)
TIAM -- A Metric for Evaluating Alignment in Text-to-Image Generation
by: Grimal, Paul, et al.
Published: (2023)
by: Grimal, Paul, et al.
Published: (2023)
Evaluation Metrics for Text Data Augmentation in NLP
by: Amadeus, Marcellus, et al.
Published: (2024)
by: Amadeus, Marcellus, et al.
Published: (2024)
MQM-Chat: Multidimensional Quality Metrics for Chat Translation
by: Li, Yunmeng, et al.
Published: (2024)
by: Li, Yunmeng, et al.
Published: (2024)
Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability
by: Hu, Xinyu, et al.
Published: (2024)
by: Hu, Xinyu, et al.
Published: (2024)
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
by: Fu, Xiao, et al.
Published: (2025)
by: Fu, Xiao, et al.
Published: (2025)
ConCISE: A Reference-Free Conciseness Evaluation Metric for LLM-Generated Answers
by: Ghafari, Seyed Mohssen, et al.
Published: (2025)
by: Ghafari, Seyed Mohssen, et al.
Published: (2025)
Towards Fine-Grained Citation Evaluation in Generated Text: A Comparative Analysis of Faithfulness Metrics
by: Zhang, Weijia, et al.
Published: (2024)
by: Zhang, Weijia, et al.
Published: (2024)
Adapting Text LLMs to Speech via Multimodal Depth Up-Scaling
by: Yano, Kazuki, et al.
Published: (2026)
by: Yano, Kazuki, et al.
Published: (2026)
Improving Explainability of Sentence-level Metrics via Edit-level Attribution for Grammatical Error Correction
by: Goto, Takumi, et al.
Published: (2024)
by: Goto, Takumi, et al.
Published: (2024)
Mitigating the Impact of Reference Quality on Evaluation of Summarization Systems with Reference-Free Metrics
by: Gigant, Théo, et al.
Published: (2024)
by: Gigant, Théo, et al.
Published: (2024)
Related Work and Citation Text Generation: A Survey
by: Li, Xiangci, et al.
Published: (2024)
by: Li, Xiangci, et al.
Published: (2024)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
by: Alam, Firoj, et al.
Published: (2026)
by: Alam, Firoj, et al.
Published: (2026)
A Judge-free LLM Open-ended Generation Benchmark Based on the Distributional Hypothesis
by: Imajo, Kentaro, et al.
Published: (2025)
by: Imajo, Kentaro, et al.
Published: (2025)
Is ChatGPT the Future of Causal Text Mining? A Comprehensive Evaluation and Analysis
by: Takayanagi, Takehiro, et al.
Published: (2024)
by: Takayanagi, Takehiro, et al.
Published: (2024)
Reference-based Metrics Disprove Themselves in Question Generation
by: Nguyen, Bang, et al.
Published: (2024)
by: Nguyen, Bang, et al.
Published: (2024)
NormEval: A Unified Multi-Metric Framework for Evaluating Semantic Fidelity in Text Normalization
by: Kafi, Md Abdullah Al, et al.
Published: (2025)
by: Kafi, Md Abdullah Al, et al.
Published: (2025)
Cobra Effect in Reference-Free Image Captioning Metrics
by: Ma, Zheng, et al.
Published: (2024)
by: Ma, Zheng, et al.
Published: (2024)
Quality and Quantity of Machine Translation References for Automatic Metrics
by: Zouhar, Vilém, et al.
Published: (2024)
by: Zouhar, Vilém, et al.
Published: (2024)
An Examination of the Robustness of Reference-Free Image Captioning Evaluation Metrics
by: Ahmadi, Saba, et al.
Published: (2023)
by: Ahmadi, Saba, et al.
Published: (2023)
CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metrics
by: Nayak, Shravan, et al.
Published: (2025)
by: Nayak, Shravan, et al.
Published: (2025)
Similar Items
-
Intrinsic Task-based Evaluation for Referring Expression Generation
by: Chen, Guanyi, et al.
Published: (2024) -
The Pitfalls of Defining Hallucination
by: van Deemter, Kees
Published: (2024) -
My Life in Artificial Intelligence: People, anecdotes, and some lessons learnt
by: van Deemter, Kees
Published: (2025) -
Computational Modelling of Plurality and Definiteness in Chinese Noun Phrases
by: Liu, Yuqi, et al.
Published: (2024) -
Textual Summarisation of Large Sets: Towards a General Approach
by: Kuptavanich, Kittipitch, et al.
Published: (2024)