A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability
Fuente:
arXiv
Salvato in:
| Autori principali: | Hu, Xinyu, Gao, Mingqi, Lin, Li, Yu, Zhenghan, Wan, Xiaojun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Analyzing and Evaluating Correlation Measures in NLG Meta-Evaluation
di: Gao, Mingqi, et al.
Pubblicazione: (2024)
di: Gao, Mingqi, et al.
Pubblicazione: (2024)
Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability
di: Hu, Xinyu, et al.
Pubblicazione: (2024)
di: Hu, Xinyu, et al.
Pubblicazione: (2024)
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
di: Chang, Jiayi, et al.
Pubblicazione: (2025)
di: Chang, Jiayi, et al.
Pubblicazione: (2025)
LLM-based NLG Evaluation: Current Status and Challenges
di: Gao, Mingqi, et al.
Pubblicazione: (2024)
di: Gao, Mingqi, et al.
Pubblicazione: (2024)
Are LLM-based Evaluators Confusing NLG Quality Criteria?
di: Hu, Xinyu, et al.
Pubblicazione: (2024)
di: Hu, Xinyu, et al.
Pubblicazione: (2024)
Better than Random: Reliable NLG Human Evaluation with Constrained Active Sampling
di: Ruan, Jie, et al.
Pubblicazione: (2024)
di: Ruan, Jie, et al.
Pubblicazione: (2024)
CFunModel: A "Funny" Language Model Capable of Chinese Humor Generation and Processing
di: Yu, Zhenghan, et al.
Pubblicazione: (2025)
di: Yu, Zhenghan, et al.
Pubblicazione: (2025)
Aspect-Guided Multi-Level Perturbation Analysis of Large Language Models in Automated Peer Review
di: Li, Jiatao, et al.
Pubblicazione: (2025)
di: Li, Jiatao, et al.
Pubblicazione: (2025)
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
di: Gao, Mingqi, et al.
Pubblicazione: (2024)
di: Gao, Mingqi, et al.
Pubblicazione: (2024)
Defining and Detecting Vulnerability in Human Evaluation Guidelines: A Preliminary Study Towards Reliable NLG Evaluation
di: Ruan, Jie, et al.
Pubblicazione: (2024)
di: Ruan, Jie, et al.
Pubblicazione: (2024)
MINOS: A Multimodal Evaluation Model for Bidirectional Generation Between Image and Text
di: Zhang, Junzhe, et al.
Pubblicazione: (2025)
di: Zhang, Junzhe, et al.
Pubblicazione: (2025)
DHP Benchmark: Are LLMs Good NLG Evaluators?
di: Wang, Yicheng, et al.
Pubblicazione: (2024)
di: Wang, Yicheng, et al.
Pubblicazione: (2024)
HAD: HAllucination Detection Language Models Based on a Comprehensive Hallucination Taxonomy
di: Xu, Fan, et al.
Pubblicazione: (2025)
di: Xu, Fan, et al.
Pubblicazione: (2025)
Large Language Models Are Active Critics in NLG Evaluation
di: Xu, Shuying, et al.
Pubblicazione: (2024)
di: Xu, Shuying, et al.
Pubblicazione: (2024)
Benchmarking Knowledge Boundary for Large Language Models: A Different Perspective on Model Evaluation
di: Yin, Xunjian, et al.
Pubblicazione: (2024)
di: Yin, Xunjian, et al.
Pubblicazione: (2024)
Evaluating Self-Generated Documents for Enhancing Retrieval-Augmented Generation with Large Language Models
di: Li, Jiatao, et al.
Pubblicazione: (2024)
di: Li, Jiatao, et al.
Pubblicazione: (2024)
What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025
di: Yang, Jing, et al.
Pubblicazione: (2026)
di: Yang, Jing, et al.
Pubblicazione: (2026)
SMART-RAG: Selection using Determinantal Matrices for Augmented Retrieval
di: Li, Jiatao, et al.
Pubblicazione: (2024)
di: Li, Jiatao, et al.
Pubblicazione: (2024)
NLG Evaluation: Past, Present, Future
di: Reiter, Ehud
Pubblicazione: (2026)
di: Reiter, Ehud
Pubblicazione: (2026)
AmbigNLG: Addressing Task Ambiguity in Instruction for NLG
di: Niwa, Ayana, et al.
Pubblicazione: (2024)
di: Niwa, Ayana, et al.
Pubblicazione: (2024)
SCOPE: Intrinsic Semantic Space Control for Mitigating Copyright Infringement in LLMs
di: Zhang, Zhenliang, et al.
Pubblicazione: (2025)
di: Zhang, Zhenliang, et al.
Pubblicazione: (2025)
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
di: Zhang, Huixuan, et al.
Pubblicazione: (2025)
di: Zhang, Huixuan, et al.
Pubblicazione: (2025)
Unveiling the Achilles' Heel of NLG Evaluators: A Unified Adversarial Framework Driven by Large Language Models
di: Chen, Yiming, et al.
Pubblicazione: (2024)
di: Chen, Yiming, et al.
Pubblicazione: (2024)
Leveraging Large Language Models for NLG Evaluation: Advances and Challenges
di: Li, Zhen, et al.
Pubblicazione: (2024)
di: Li, Zhen, et al.
Pubblicazione: (2024)
Is Reference Necessary in the Evaluation of NLG Systems? When and Where?
di: Sheng, Shuqian, et al.
Pubblicazione: (2024)
di: Sheng, Shuqian, et al.
Pubblicazione: (2024)
TailNLG: A Multilingual Benchmark Addressing Verbalization of Long-Tail Entities
di: Draetta, Lia, et al.
Pubblicazione: (2026)
di: Draetta, Lia, et al.
Pubblicazione: (2026)
How to Select Datapoints for Efficient Human Evaluation of NLG Models?
di: Zouhar, Vilém, et al.
Pubblicazione: (2025)
di: Zouhar, Vilém, et al.
Pubblicazione: (2025)
Not All Metrics Are Guilty: Improving NLG Evaluation by Diversifying References
di: Tang, Tianyi, et al.
Pubblicazione: (2023)
di: Tang, Tianyi, et al.
Pubblicazione: (2023)
Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts
di: Hong, Hanhua, et al.
Pubblicazione: (2025)
di: Hong, Hanhua, et al.
Pubblicazione: (2025)
OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs
di: Kartáč, Ivan, et al.
Pubblicazione: (2025)
di: Kartáč, Ivan, et al.
Pubblicazione: (2025)
MCQA-Eval: Efficient Confidence Evaluation in NLG with Gold-Standard Correctness Labels
di: Liu, Xiaoou, et al.
Pubblicazione: (2025)
di: Liu, Xiaoou, et al.
Pubblicazione: (2025)
Error-Robust Retrieval for Chinese Spelling Check
di: Yin, Xunjian, et al.
Pubblicazione: (2022)
di: Yin, Xunjian, et al.
Pubblicazione: (2022)
KBE-DME: Dynamic Multimodal Evaluation via Knowledge Enhanced Benchmark Evolution
di: Zhang, Junzhe, et al.
Pubblicazione: (2025)
di: Zhang, Junzhe, et al.
Pubblicazione: (2025)
LLMSR@XLLM25: An Empirical Study of LLM for Structural Reasoning
di: Li, Xinye, et al.
Pubblicazione: (2025)
di: Li, Xinye, et al.
Pubblicazione: (2025)
MC-MKE: A Fine-Grained Multimodal Knowledge Editing Benchmark Emphasizing Modality Consistency
di: Zhang, Junzhe, et al.
Pubblicazione: (2024)
di: Zhang, Junzhe, et al.
Pubblicazione: (2024)
Integration of LLM Quality Assurance into an NLG System
di: Chen, Ching-Yi, et al.
Pubblicazione: (2025)
di: Chen, Ching-Yi, et al.
Pubblicazione: (2025)
A Dual-Perspective Metaphor Detection Framework Using Large Language Models
di: Lin, Yujie, et al.
Pubblicazione: (2024)
di: Lin, Yujie, et al.
Pubblicazione: (2024)
UniAIDet: A Unified and Universal Benchmark for AI-Generated Image Content Detection and Localization
di: Zhang, Huixuan, et al.
Pubblicazione: (2025)
di: Zhang, Huixuan, et al.
Pubblicazione: (2025)
PaCoST: Paired Confidence Significance Testing for Benchmark Contamination Detection in Large Language Models
di: Zhang, Huixuan, et al.
Pubblicazione: (2024)
di: Zhang, Huixuan, et al.
Pubblicazione: (2024)
Rethinking Meeting Effectiveness: A Benchmark and Framework for Temporal Fine-grained Automatic Meeting Effectiveness Evaluation
di: Li, Yihang, et al.
Pubblicazione: (2026)
di: Li, Yihang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Analyzing and Evaluating Correlation Measures in NLG Meta-Evaluation
di: Gao, Mingqi, et al.
Pubblicazione: (2024) -
Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability
di: Hu, Xinyu, et al.
Pubblicazione: (2024) -
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
di: Chang, Jiayi, et al.
Pubblicazione: (2025) -
LLM-based NLG Evaluation: Current Status and Challenges
di: Gao, Mingqi, et al.
Pubblicazione: (2024) -
Are LLM-based Evaluators Confusing NLG Quality Criteria?
di: Hu, Xinyu, et al.
Pubblicazione: (2024)