Not All Metrics Are Guilty: Improving NLG Evaluation by Diversifying References
Fuente:
arXiv
Guardado en:
| Autores principales: | Tang, Tianyi, Lu, Hongyuan, Jiang, Yuchen Eleanor, Huang, Haoyang, Zhang, Dongdong, Zhao, Wayne Xin, Kocmi, Tom, Wei, Furu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Chain-of-Dictionary Prompting Elicits Translation in Large Language Models
por: Lu, Hongyuan, et al.
Publicado: (2023)
por: Lu, Hongyuan, et al.
Publicado: (2023)
SLIDE: Reference-free Evaluation for Machine Translation using a Sliding Document Window
por: Raunak, Vikas, et al.
Publicado: (2023)
por: Raunak, Vikas, et al.
Publicado: (2023)
Pearmut: Human Evaluation of Translation Made Trivial
por: Zouhar, Vilém, et al.
Publicado: (2026)
por: Zouhar, Vilém, et al.
Publicado: (2026)
Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models
por: Tang, Tianyi, et al.
Publicado: (2024)
por: Tang, Tianyi, et al.
Publicado: (2024)
Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies
por: Kocmi, Tom, et al.
Publicado: (2024)
por: Kocmi, Tom, et al.
Publicado: (2024)
Is Reference Necessary in the Evaluation of NLG Systems? When and Where?
por: Sheng, Shuqian, et al.
Publicado: (2024)
por: Sheng, Shuqian, et al.
Publicado: (2024)
AI-Assisted Human Evaluation of Machine Translation
por: Zouhar, Vilém, et al.
Publicado: (2024)
por: Zouhar, Vilém, et al.
Publicado: (2024)
OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs
por: Kartáč, Ivan, et al.
Publicado: (2025)
por: Kartáč, Ivan, et al.
Publicado: (2025)
LAP for "Growing Up Guilty."
por: Schwartz, Sheila
Publicado: (1985)
por: Schwartz, Sheila
Publicado: (1985)
Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability
por: Hu, Xinyu, et al.
Publicado: (2024)
por: Hu, Xinyu, et al.
Publicado: (2024)
Large Language Models Are Active Critics in NLG Evaluation
por: Xu, Shuying, et al.
Publicado: (2024)
por: Xu, Shuying, et al.
Publicado: (2024)
Guilty Land / Patrick van Rensburg
por: Rensburg, Patrick van
Publicado: (1962)
por: Rensburg, Patrick van
Publicado: (1962)
Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation
por: Kreutzer, Julia, et al.
Publicado: (2025)
por: Kreutzer, Julia, et al.
Publicado: (2025)
NLG Evaluation: Past, Present, Future
por: Reiter, Ehud
Publicado: (2026)
por: Reiter, Ehud
Publicado: (2026)
Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models
por: Lu, Qingyu, et al.
Publicado: (2023)
por: Lu, Qingyu, et al.
Publicado: (2023)
Optimal Decision Mechanisms for Committees: Acquitting the Guilty
por: Kattwinkel, Deniz, et al.
Publicado: (2024)
por: Kattwinkel, Deniz, et al.
Publicado: (2024)
AmbigNLG: Addressing Task Ambiguity in Instruction for NLG
por: Niwa, Ayana, et al.
Publicado: (2024)
por: Niwa, Ayana, et al.
Publicado: (2024)
A Privacy-Preserving Data Collection Method for Diversified Statistical Analysis
por: Jiang, Hao, et al.
Publicado: (2025)
por: Jiang, Hao, et al.
Publicado: (2025)
Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts
por: Hong, Hanhua, et al.
Publicado: (2025)
por: Hong, Hanhua, et al.
Publicado: (2025)
Analyzing and Evaluating Correlation Measures in NLG Meta-Evaluation
por: Gao, Mingqi, et al.
Publicado: (2024)
por: Gao, Mingqi, et al.
Publicado: (2024)
DHP Benchmark: Are LLMs Good NLG Evaluators?
por: Wang, Yicheng, et al.
Publicado: (2024)
por: Wang, Yicheng, et al.
Publicado: (2024)
SAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text
por: Ghosh, Reshmi, et al.
Publicado: (2024)
por: Ghosh, Reshmi, et al.
Publicado: (2024)
New Reference: Diversifying Service Delivery.
por: Garner, Imogen
Publicado: (1999)
por: Garner, Imogen
Publicado: (1999)
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
por: Chang, Jiayi, et al.
Publicado: (2025)
por: Chang, Jiayi, et al.
Publicado: (2025)
BAMBOO: A Comprehensive Benchmark for Evaluating Long Text Modeling Capacities of Large Language Models
por: Dong, Zican, et al.
Publicado: (2023)
por: Dong, Zican, et al.
Publicado: (2023)
Are LLM-based Evaluators Confusing NLG Quality Criteria?
por: Hu, Xinyu, et al.
Publicado: (2024)
por: Hu, Xinyu, et al.
Publicado: (2024)
LLM-based NLG Evaluation: Current Status and Challenges
por: Gao, Mingqi, et al.
Publicado: (2024)
por: Gao, Mingqi, et al.
Publicado: (2024)
A Survey on Long Text Modeling with Transformers
por: Dong, Zican, et al.
Publicado: (2023)
por: Dong, Zican, et al.
Publicado: (2023)
The Copyright Caper. Transferring Records to Tape: How Guilty Are You?
por: Shaw, Jim
Publicado: (1973)
por: Shaw, Jim
Publicado: (1973)
TIGERScore: Towards Building Explainable Metric for All Text Generation Tasks
por: Jiang, Dongfu, et al.
Publicado: (2023)
por: Jiang, Dongfu, et al.
Publicado: (2023)
VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
por: Jiang, Lingjie, et al.
Publicado: (2025)
por: Jiang, Lingjie, et al.
Publicado: (2025)
MCQA-Eval: Efficient Confidence Evaluation in NLG with Gold-Standard Correctness Labels
por: Liu, Xiaoou, et al.
Publicado: (2025)
por: Liu, Xiaoou, et al.
Publicado: (2025)
How to Select Datapoints for Efficient Human Evaluation of NLG Models?
por: Zouhar, Vilém, et al.
Publicado: (2025)
por: Zouhar, Vilém, et al.
Publicado: (2025)
Leveraging Large Language Models for NLG Evaluation: Advances and Challenges
por: Li, Zhen, et al.
Publicado: (2024)
por: Li, Zhen, et al.
Publicado: (2024)
InstructDET: Diversifying Referring Object Detection with Generalized Instructions
por: Dang, Ronghao, et al.
Publicado: (2023)
por: Dang, Ronghao, et al.
Publicado: (2023)
The statistical advantage of automatic NLG metrics at the system level
por: Wei, Johnny Tian-Zheng, et al.
Publicado: (2021)
por: Wei, Johnny Tian-Zheng, et al.
Publicado: (2021)
Reasoning with Exploration: An Entropy Perspective
por: Cheng, Daixuan, et al.
Publicado: (2025)
por: Cheng, Daixuan, et al.
Publicado: (2025)
From Abstract to Contextual: What LLMs Still Cannot Do in Mathematics
por: Cao, Bowen, et al.
Publicado: (2026)
por: Cao, Bowen, et al.
Publicado: (2026)
Direct-Scoring NLG Evaluators Can Use Pairwise Comparisons Too
por: Lawrence, Logan, et al.
Publicado: (2025)
por: Lawrence, Logan, et al.
Publicado: (2025)
Guilty Plea Appeals Across Two States: Who Appeals and Who Wins?
por: Jacqueline G. Lee, et al.
Publicado: (2026)
por: Jacqueline G. Lee, et al.
Publicado: (2026)
Ejemplares similares
-
Chain-of-Dictionary Prompting Elicits Translation in Large Language Models
por: Lu, Hongyuan, et al.
Publicado: (2023) -
SLIDE: Reference-free Evaluation for Machine Translation using a Sliding Document Window
por: Raunak, Vikas, et al.
Publicado: (2023) -
Pearmut: Human Evaluation of Translation Made Trivial
por: Zouhar, Vilém, et al.
Publicado: (2026) -
Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models
por: Tang, Tianyi, et al.
Publicado: (2024) -
Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies
por: Kocmi, Tom, et al.
Publicado: (2024)