Gespeichert in:
| Hauptverfasser: | Gao, Mingqi, Hu, Xinyu, Lin, Li, Wan, Xiaojun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2410.16834 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability
von: Hu, Xinyu, et al.
Veröffentlicht: (2025)
von: Hu, Xinyu, et al.
Veröffentlicht: (2025)
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
von: Chang, Jiayi, et al.
Veröffentlicht: (2025)
von: Chang, Jiayi, et al.
Veröffentlicht: (2025)
Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability
von: Hu, Xinyu, et al.
Veröffentlicht: (2024)
von: Hu, Xinyu, et al.
Veröffentlicht: (2024)
LLM-based NLG Evaluation: Current Status and Challenges
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
Are LLM-based Evaluators Confusing NLG Quality Criteria?
von: Hu, Xinyu, et al.
Veröffentlicht: (2024)
von: Hu, Xinyu, et al.
Veröffentlicht: (2024)
Better than Random: Reliable NLG Human Evaluation with Constrained Active Sampling
von: Ruan, Jie, et al.
Veröffentlicht: (2024)
von: Ruan, Jie, et al.
Veröffentlicht: (2024)
MINOS: A Multimodal Evaluation Model for Bidirectional Generation Between Image and Text
von: Zhang, Junzhe, et al.
Veröffentlicht: (2025)
von: Zhang, Junzhe, et al.
Veröffentlicht: (2025)
Defining and Detecting Vulnerability in Human Evaluation Guidelines: A Preliminary Study Towards Reliable NLG Evaluation
von: Ruan, Jie, et al.
Veröffentlicht: (2024)
von: Ruan, Jie, et al.
Veröffentlicht: (2024)
Aspect-Guided Multi-Level Perturbation Analysis of Large Language Models in Automated Peer Review
von: Li, Jiatao, et al.
Veröffentlicht: (2025)
von: Li, Jiatao, et al.
Veröffentlicht: (2025)
What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025
von: Yang, Jing, et al.
Veröffentlicht: (2026)
von: Yang, Jing, et al.
Veröffentlicht: (2026)
Large Language Models Are Active Critics in NLG Evaluation
von: Xu, Shuying, et al.
Veröffentlicht: (2024)
von: Xu, Shuying, et al.
Veröffentlicht: (2024)
Evaluating Self-Generated Documents for Enhancing Retrieval-Augmented Generation with Large Language Models
von: Li, Jiatao, et al.
Veröffentlicht: (2024)
von: Li, Jiatao, et al.
Veröffentlicht: (2024)
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
NLG Evaluation: Past, Present, Future
von: Reiter, Ehud
Veröffentlicht: (2026)
von: Reiter, Ehud
Veröffentlicht: (2026)
DHP Benchmark: Are LLMs Good NLG Evaluators?
von: Wang, Yicheng, et al.
Veröffentlicht: (2024)
von: Wang, Yicheng, et al.
Veröffentlicht: (2024)
Is Reference Necessary in the Evaluation of NLG Systems? When and Where?
von: Sheng, Shuqian, et al.
Veröffentlicht: (2024)
von: Sheng, Shuqian, et al.
Veröffentlicht: (2024)
Leveraging Large Language Models for NLG Evaluation: Advances and Challenges
von: Li, Zhen, et al.
Veröffentlicht: (2024)
von: Li, Zhen, et al.
Veröffentlicht: (2024)
SMART-RAG: Selection using Determinantal Matrices for Augmented Retrieval
von: Li, Jiatao, et al.
Veröffentlicht: (2024)
von: Li, Jiatao, et al.
Veröffentlicht: (2024)
How to Select Datapoints for Efficient Human Evaluation of NLG Models?
von: Zouhar, Vilém, et al.
Veröffentlicht: (2025)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2025)
Not All Metrics Are Guilty: Improving NLG Evaluation by Diversifying References
von: Tang, Tianyi, et al.
Veröffentlicht: (2023)
von: Tang, Tianyi, et al.
Veröffentlicht: (2023)
MCQA-Eval: Efficient Confidence Evaluation in NLG with Gold-Standard Correctness Labels
von: Liu, Xiaoou, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoou, et al.
Veröffentlicht: (2025)
AmbigNLG: Addressing Task Ambiguity in Instruction for NLG
von: Niwa, Ayana, et al.
Veröffentlicht: (2024)
von: Niwa, Ayana, et al.
Veröffentlicht: (2024)
OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs
von: Kartáč, Ivan, et al.
Veröffentlicht: (2025)
von: Kartáč, Ivan, et al.
Veröffentlicht: (2025)
SCOPE: Intrinsic Semantic Space Control for Mitigating Copyright Infringement in LLMs
von: Zhang, Zhenliang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenliang, et al.
Veröffentlicht: (2025)
CFunModel: A "Funny" Language Model Capable of Chinese Humor Generation and Processing
von: Yu, Zhenghan, et al.
Veröffentlicht: (2025)
von: Yu, Zhenghan, et al.
Veröffentlicht: (2025)
Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts
von: Hong, Hanhua, et al.
Veröffentlicht: (2025)
von: Hong, Hanhua, et al.
Veröffentlicht: (2025)
Evaluating, Understanding, and Improving Constrained Text Generation for Large Language Models
von: Chen, Xiang, et al.
Veröffentlicht: (2023)
von: Chen, Xiang, et al.
Veröffentlicht: (2023)
Direct-Scoring NLG Evaluators Can Use Pairwise Comparisons Too
von: Lawrence, Logan, et al.
Veröffentlicht: (2025)
von: Lawrence, Logan, et al.
Veröffentlicht: (2025)
Unveiling the Achilles' Heel of NLG Evaluators: A Unified Adversarial Framework Driven by Large Language Models
von: Chen, Yiming, et al.
Veröffentlicht: (2024)
von: Chen, Yiming, et al.
Veröffentlicht: (2024)
Analyzing Cognitive Differences Among Large Language Models through the Lens of Social Worldview
von: Li, Jiatao, et al.
Veröffentlicht: (2025)
von: Li, Jiatao, et al.
Veröffentlicht: (2025)
Error-Robust Retrieval for Chinese Spelling Check
von: Yin, Xunjian, et al.
Veröffentlicht: (2022)
von: Yin, Xunjian, et al.
Veröffentlicht: (2022)
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
LLMSR@XLLM25: An Empirical Study of LLM for Structural Reasoning
von: Li, Xinye, et al.
Veröffentlicht: (2025)
von: Li, Xinye, et al.
Veröffentlicht: (2025)
SAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text
von: Ghosh, Reshmi, et al.
Veröffentlicht: (2024)
von: Ghosh, Reshmi, et al.
Veröffentlicht: (2024)
LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
von: Liusie, Adian, et al.
Veröffentlicht: (2023)
von: Liusie, Adian, et al.
Veröffentlicht: (2023)
KBE-DME: Dynamic Multimodal Evaluation via Knowledge Enhanced Benchmark Evolution
von: Zhang, Junzhe, et al.
Veröffentlicht: (2025)
von: Zhang, Junzhe, et al.
Veröffentlicht: (2025)
Integration of LLM Quality Assurance into an NLG System
von: Chen, Ching-Yi, et al.
Veröffentlicht: (2025)
von: Chen, Ching-Yi, et al.
Veröffentlicht: (2025)
Benchmarking Knowledge Boundary for Large Language Models: A Different Perspective on Model Evaluation
von: Yin, Xunjian, et al.
Veröffentlicht: (2024)
von: Yin, Xunjian, et al.
Veröffentlicht: (2024)
Exploring Causal Effect of Social Bias on Faithfulness Hallucinations in Large Language Models
von: Zhang, Zhenliang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenliang, et al.
Veröffentlicht: (2025)
The statistical advantage of automatic NLG metrics at the system level
von: Wei, Johnny Tian-Zheng, et al.
Veröffentlicht: (2021)
von: Wei, Johnny Tian-Zheng, et al.
Veröffentlicht: (2021)
Ähnliche Einträge
-
A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability
von: Hu, Xinyu, et al.
Veröffentlicht: (2025) -
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
von: Chang, Jiayi, et al.
Veröffentlicht: (2025) -
Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability
von: Hu, Xinyu, et al.
Veröffentlicht: (2024) -
LLM-based NLG Evaluation: Current Status and Challenges
von: Gao, Mingqi, et al.
Veröffentlicht: (2024) -
Are LLM-based Evaluators Confusing NLG Quality Criteria?
von: Hu, Xinyu, et al.
Veröffentlicht: (2024)