Automatic Metrics in Natural Language Generation: A Survey of Current Evaluation Practices
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Schmidtová, Patrícia, Mahamood, Saad, Balloccu, Simone, Dušek, Ondřej, Gatt, Albert, Gkatzia, Dimitra, Howcroft, David M., Plátek, Ondřej, Sivaprasad, Adarsa |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
factgenie: A Framework for Span-based Evaluation of Generated Texts
von: Kasner, Zdeněk, et al.
Veröffentlicht: (2024)
von: Kasner, Zdeněk, et al.
Veröffentlicht: (2024)
LLMs as Span Annotators: A Comparative Study of LLMs and Humans
von: Kasner, Zdeněk, et al.
Veröffentlicht: (2025)
von: Kasner, Zdeněk, et al.
Veröffentlicht: (2025)
Real-World Summarization: When Evaluation Reaches Its Limits
von: Schmidtová, Patrícia, et al.
Veröffentlicht: (2025)
von: Schmidtová, Patrícia, et al.
Veröffentlicht: (2025)
Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs
von: Balloccu, Simone, et al.
Veröffentlicht: (2024)
von: Balloccu, Simone, et al.
Veröffentlicht: (2024)
FreshTab: Sourcing Fresh Data for Table-to-Text Generation Evaluation
von: Onderková, Kristýna, et al.
Veröffentlicht: (2025)
von: Onderková, Kristýna, et al.
Veröffentlicht: (2025)
When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition
von: Li, Karen Jia-Hui, et al.
Veröffentlicht: (2025)
von: Li, Karen Jia-Hui, et al.
Veröffentlicht: (2025)
OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs
von: Kartáč, Ivan, et al.
Veröffentlicht: (2025)
von: Kartáč, Ivan, et al.
Veröffentlicht: (2025)
Linguistically Communicating Uncertainty in Patient-Facing Risk Prediction Models
von: Sivaprasad, Adarsa, et al.
Veröffentlicht: (2024)
von: Sivaprasad, Adarsa, et al.
Veröffentlicht: (2024)
Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach
von: Wojciechowski, Adam, et al.
Veröffentlicht: (2024)
von: Wojciechowski, Adam, et al.
Veröffentlicht: (2024)
Beyond Traditional Benchmarks: Analyzing Behaviors of Open LLMs on Data-to-Text Generation
von: Kasner, Zdeněk, et al.
Veröffentlicht: (2024)
von: Kasner, Zdeněk, et al.
Veröffentlicht: (2024)
LLM Agents Implement an NLG System from Scratch: Building Interpretable Rule-Based RDF-to-Text Generators
von: Lango, Mateusz, et al.
Veröffentlicht: (2025)
von: Lango, Mateusz, et al.
Veröffentlicht: (2025)
Text Style Transfer: An Introductory Overview
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2024)
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2024)
LEEETs-Dial: Linguistic Entrainment in End-to-End Task-oriented Dialogue systems
von: Kumar, Nalin, et al.
Veröffentlicht: (2023)
von: Kumar, Nalin, et al.
Veröffentlicht: (2023)
AnimatedLLM: Explaining LLMs with Interactive Visualizations
von: Kasner, Zdeněk, et al.
Veröffentlicht: (2025)
von: Kasner, Zdeněk, et al.
Veröffentlicht: (2025)
Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics?
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2025)
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2025)
Strategies for Span Labeling with Large Language Models
von: Semin, Danil, et al.
Veröffentlicht: (2026)
von: Semin, Danil, et al.
Veröffentlicht: (2026)
Evaluation of Human-Understandability of Global Model Explanations using Decision Tree
von: Sivaprasad, Adarsa, et al.
Veröffentlicht: (2023)
von: Sivaprasad, Adarsa, et al.
Veröffentlicht: (2023)
Leveraging Large Language Models for Building Interpretable Rule-Based Data-to-Text Systems
von: Warczyński, Jędrzej, et al.
Veröffentlicht: (2025)
von: Warczyński, Jędrzej, et al.
Veröffentlicht: (2025)
Quality and Quantity of Machine Translation References for Automatic Metrics
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
A Survey of Text Style Transfer: Applications and Ethical Implications
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2024)
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2024)
Are Large Language Models Actually Good at Text Style Transfer?
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2024)
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2024)
Dark Side Augmentation: Generating Diverse Night Examples for Metric Learning
von: Mohwald, Albert, et al.
Veröffentlicht: (2023)
von: Mohwald, Albert, et al.
Veröffentlicht: (2023)
SRS-Stories: Vocabulary-constrained multilingual story generation for language learning
von: Kamzela, Wiktor, et al.
Veröffentlicht: (2025)
von: Kamzela, Wiktor, et al.
Veröffentlicht: (2025)
Reasoning Gets Harder for LLMs Inside A Dialogue
von: Kartáč, Ivan, et al.
Veröffentlicht: (2026)
von: Kartáč, Ivan, et al.
Veröffentlicht: (2026)
Sentence Embeddings as an intermediate target in end-to-end summarisation
von: Zembrzuski, Maciej, et al.
Veröffentlicht: (2025)
von: Zembrzuski, Maciej, et al.
Veröffentlicht: (2025)
A Practical Specification Language for Automatic Quantum Program Verification (Technical Report)
von: Tsai, Wei-Lun, et al.
Veröffentlicht: (2026)
von: Tsai, Wei-Lun, et al.
Veröffentlicht: (2026)
Neural Networks Based Domain Adaptation in Spectroscopic Sky Surveys
von: Podsztavek, Ondřej
Veröffentlicht: (2020)
von: Podsztavek, Ondřej
Veröffentlicht: (2020)
Proof of the Generalization of the Sawayama-Thébault Theorem
von: Płatek, Miłosz
Veröffentlicht: (2026)
von: Płatek, Miłosz
Veröffentlicht: (2026)
The Scaffold Effect: How Prompt Framing Drives Apparent Multimodal Gains in Clinical VLM Evaluation
von: Vu, Doan Nam Long, et al.
Veröffentlicht: (2026)
von: Vu, Doan Nam Long, et al.
Veröffentlicht: (2026)
Morphological Analysis for the Maltese Language: The Challenges of a Hybrid System
von: Borg, Claudia, et al.
Veröffentlicht: (2017)
von: Borg, Claudia, et al.
Veröffentlicht: (2017)
Evaluating LLM-Generated Versus Human-Authored Responses in Role-Play Dialogues
von: Lu, Dongxu, et al.
Veröffentlicht: (2025)
von: Lu, Dongxu, et al.
Veröffentlicht: (2025)
How Do People Quantify Naturally: Evidence from Mandarin Picture Description
von: Zhang, Yayun, et al.
Veröffentlicht: (2026)
von: Zhang, Yayun, et al.
Veröffentlicht: (2026)
End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
von: Luu, Nam, et al.
Veröffentlicht: (2025)
von: Luu, Nam, et al.
Veröffentlicht: (2025)
Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding
von: Ignatev, Daniil, et al.
Veröffentlicht: (2025)
von: Ignatev, Daniil, et al.
Veröffentlicht: (2025)
A Survey of Automatic Hallucination Evaluation on Natural Language Generation
von: Qi, Siya, et al.
Veröffentlicht: (2024)
von: Qi, Siya, et al.
Veröffentlicht: (2024)
Patient-Centred Explainability in IVF Outcome Prediction
von: Sivaprasad, Adarsa, et al.
Veröffentlicht: (2025)
von: Sivaprasad, Adarsa, et al.
Veröffentlicht: (2025)
Calculation of a force effect from muscle action to a quaternion-based musculoskeletal model
von: Zoufaly, Ondrej, et al.
Veröffentlicht: (2025)
von: Zoufaly, Ondrej, et al.
Veröffentlicht: (2025)
Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation
von: Polák, Peter, et al.
Veröffentlicht: (2025)
von: Polák, Peter, et al.
Veröffentlicht: (2025)
Subjective Code Preferences in Experts and Large Language Models
von: Mokhova, Anna, et al.
Veröffentlicht: (2026)
von: Mokhova, Anna, et al.
Veröffentlicht: (2026)
Intrinsic vs. Extrinsic Evaluation of Czech Sentence Embeddings: Semantic Relevance Doesn't Help with MT Evaluation
von: Barančíková, Petra, et al.
Veröffentlicht: (2025)
von: Barančíková, Petra, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
factgenie: A Framework for Span-based Evaluation of Generated Texts
von: Kasner, Zdeněk, et al.
Veröffentlicht: (2024) -
LLMs as Span Annotators: A Comparative Study of LLMs and Humans
von: Kasner, Zdeněk, et al.
Veröffentlicht: (2025) -
Real-World Summarization: When Evaluation Reaches Its Limits
von: Schmidtová, Patrícia, et al.
Veröffentlicht: (2025) -
Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs
von: Balloccu, Simone, et al.
Veröffentlicht: (2024) -
FreshTab: Sourcing Fresh Data for Table-to-Text Generation Evaluation
von: Onderková, Kristýna, et al.
Veröffentlicht: (2025)