Integration of LLM Quality Assurance into an NLG System
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Ching-Yi, Heininger, Johanna, Schneider, Adela, Eckard, Christian, Madsack, Andreas, Weißgraeber, Robert |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluation of NMT-Assisted Grammar Transfer for a Multi-Language Configurable Data-to-Text System
by: Madsack, Andreas, et al.
Published: (2025)
by: Madsack, Andreas, et al.
Published: (2025)
Are LLM-based Evaluators Confusing NLG Quality Criteria?
by: Hu, Xinyu, et al.
Published: (2024)
by: Hu, Xinyu, et al.
Published: (2024)
Is Reference Necessary in the Evaluation of NLG Systems? When and Where?
by: Sheng, Shuqian, et al.
Published: (2024)
by: Sheng, Shuqian, et al.
Published: (2024)
AmbigNLG: Addressing Task Ambiguity in Instruction for NLG
by: Niwa, Ayana, et al.
Published: (2024)
by: Niwa, Ayana, et al.
Published: (2024)
LLM-based NLG Evaluation: Current Status and Challenges
by: Gao, Mingqi, et al.
Published: (2024)
by: Gao, Mingqi, et al.
Published: (2024)
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
by: Chang, Jiayi, et al.
Published: (2025)
by: Chang, Jiayi, et al.
Published: (2025)
Counterfactual Fairness Evaluation of LLM-Based Contact Center Agent Quality Assurance System
by: Mayilvaghanan, Kawin, et al.
Published: (2026)
by: Mayilvaghanan, Kawin, et al.
Published: (2026)
NLG Evaluation: Past, Present, Future
by: Reiter, Ehud
Published: (2026)
by: Reiter, Ehud
Published: (2026)
Leveraging Large Language Models for NLG Evaluation: Advances and Challenges
by: Li, Zhen, et al.
Published: (2024)
by: Li, Zhen, et al.
Published: (2024)
LLM Agents Implement an NLG System from Scratch: Building Interpretable Rule-Based RDF-to-Text Generators
by: Lango, Mateusz, et al.
Published: (2025)
by: Lango, Mateusz, et al.
Published: (2025)
Large Language Models Are Active Critics in NLG Evaluation
by: Xu, Shuying, et al.
Published: (2024)
by: Xu, Shuying, et al.
Published: (2024)
The statistical advantage of automatic NLG metrics at the system level
by: Wei, Johnny Tian-Zheng, et al.
Published: (2021)
by: Wei, Johnny Tian-Zheng, et al.
Published: (2021)
Analyzing and Evaluating Correlation Measures in NLG Meta-Evaluation
by: Gao, Mingqi, et al.
Published: (2024)
by: Gao, Mingqi, et al.
Published: (2024)
DHP Benchmark: Are LLMs Good NLG Evaluators?
by: Wang, Yicheng, et al.
Published: (2024)
by: Wang, Yicheng, et al.
Published: (2024)
LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
by: Liusie, Adian, et al.
Published: (2023)
by: Liusie, Adian, et al.
Published: (2023)
How to Select Datapoints for Efficient Human Evaluation of NLG Models?
by: Zouhar, Vilém, et al.
Published: (2025)
by: Zouhar, Vilém, et al.
Published: (2025)
Not All Metrics Are Guilty: Improving NLG Evaluation by Diversifying References
by: Tang, Tianyi, et al.
Published: (2023)
by: Tang, Tianyi, et al.
Published: (2023)
IFDID: Information Filter upon Diversity-Improved Decoding for Diversity-Faithfulness Tradeoff in NLG
by: Meng, Han, et al.
Published: (2022)
by: Meng, Han, et al.
Published: (2022)
OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs
by: Kartáč, Ivan, et al.
Published: (2025)
by: Kartáč, Ivan, et al.
Published: (2025)
From Instruction to Output: The Role of Prompting in Modern NLG
by: Zaib, Munazza, et al.
Published: (2026)
by: Zaib, Munazza, et al.
Published: (2026)
An End-to-End System for Culturally-Attuned Driving Feedback using a Dual-Component NLG Engine
by: Thompson, Iniakpokeikiye Peter, et al.
Published: (2025)
by: Thompson, Iniakpokeikiye Peter, et al.
Published: (2025)
Unveiling the Achilles' Heel of NLG Evaluators: A Unified Adversarial Framework Driven by Large Language Models
by: Chen, Yiming, et al.
Published: (2024)
by: Chen, Yiming, et al.
Published: (2024)
A Systematic Review of Data-to-Text NLG
by: Osuji, Chinonso Cynthia, et al.
Published: (2024)
by: Osuji, Chinonso Cynthia, et al.
Published: (2024)
Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability
by: Hu, Xinyu, et al.
Published: (2024)
by: Hu, Xinyu, et al.
Published: (2024)
TailNLG: A Multilingual Benchmark Addressing Verbalization of Long-Tail Entities
by: Draetta, Lia, et al.
Published: (2026)
by: Draetta, Lia, et al.
Published: (2026)
MCQA-Eval: Efficient Confidence Evaluation in NLG with Gold-Standard Correctness Labels
by: Liu, Xiaoou, et al.
Published: (2025)
by: Liu, Xiaoou, et al.
Published: (2025)
"One-Size-Fits-All"? Examining Expectations around What Constitute "Fair" or "Good" NLG System Behaviors
by: Lucy, Li, et al.
Published: (2023)
by: Lucy, Li, et al.
Published: (2023)
Stability Analysis of ChatGPT-based Sentiment Analysis in AI Quality Assurance
by: Ouyang, Tinghui, et al.
Published: (2024)
by: Ouyang, Tinghui, et al.
Published: (2024)
A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability
by: Hu, Xinyu, et al.
Published: (2025)
by: Hu, Xinyu, et al.
Published: (2025)
What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025
by: Yang, Jing, et al.
Published: (2026)
by: Yang, Jing, et al.
Published: (2026)
SAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text
by: Ghosh, Reshmi, et al.
Published: (2024)
by: Ghosh, Reshmi, et al.
Published: (2024)
Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts
by: Hong, Hanhua, et al.
Published: (2025)
by: Hong, Hanhua, et al.
Published: (2025)
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
by: Baumgärtner, Tim, et al.
Published: (2026)
by: Baumgärtner, Tim, et al.
Published: (2026)
Better than Random: Reliable NLG Human Evaluation with Constrained Active Sampling
by: Ruan, Jie, et al.
Published: (2024)
by: Ruan, Jie, et al.
Published: (2024)
Weakly Supervised Fine-grained Span-Level Framework for Chinese Radiology Report Quality Assurance
by: Wang, Kaiyu, et al.
Published: (2025)
by: Wang, Kaiyu, et al.
Published: (2025)
vLLM Hook v0: A Plug-in for Programming Model Internals on vLLM
by: Ko, Ching-Yun, et al.
Published: (2026)
by: Ko, Ching-Yun, et al.
Published: (2026)
Advancing Risk and Quality Assurance: A RAG Chatbot for Improved Regulatory Compliance
by: Hillebrand, Lars, et al.
Published: (2025)
by: Hillebrand, Lars, et al.
Published: (2025)
Direct-Scoring NLG Evaluators Can Use Pairwise Comparisons Too
by: Lawrence, Logan, et al.
Published: (2025)
by: Lawrence, Logan, et al.
Published: (2025)
Diagnosing Translated Benchmarks: An Automated Quality Assurance Study of the EU20 Benchmark Suite
by: Thellmann, Klaudia, et al.
Published: (2026)
by: Thellmann, Klaudia, et al.
Published: (2026)
Defining and Detecting Vulnerability in Human Evaluation Guidelines: A Preliminary Study Towards Reliable NLG Evaluation
by: Ruan, Jie, et al.
Published: (2024)
by: Ruan, Jie, et al.
Published: (2024)
Similar Items
-
Evaluation of NMT-Assisted Grammar Transfer for a Multi-Language Configurable Data-to-Text System
by: Madsack, Andreas, et al.
Published: (2025) -
Are LLM-based Evaluators Confusing NLG Quality Criteria?
by: Hu, Xinyu, et al.
Published: (2024) -
Is Reference Necessary in the Evaluation of NLG Systems? When and Where?
by: Sheng, Shuqian, et al.
Published: (2024) -
AmbigNLG: Addressing Task Ambiguity in Instruction for NLG
by: Niwa, Ayana, et al.
Published: (2024) -
LLM-based NLG Evaluation: Current Status and Challenges
by: Gao, Mingqi, et al.
Published: (2024)