DHP Benchmark: Are LLMs Good NLG Evaluators?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yicheng, Yuan, Jiayi, Chuang, Yu-Neng, Wang, Zhuoer, Liu, Yingchi, Cusick, Mark, Kulkarni, Param, Ji, Zhengping, Ibrahim, Yasser, Hu, Xia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FairLENS: Assessing Fairness in Law Enforcement Speech Recognition
von: Wang, Yicheng, et al.
Veröffentlicht: (2024)
von: Wang, Yicheng, et al.
Veröffentlicht: (2024)
Auto-Drafting Police Reports from Noisy ASR Outputs: A Trust-Centered LLM Approach
von: Kulkarni, Param, et al.
Veröffentlicht: (2025)
von: Kulkarni, Param, et al.
Veröffentlicht: (2025)
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
von: Chang, Jiayi, et al.
Veröffentlicht: (2025)
von: Chang, Jiayi, et al.
Veröffentlicht: (2025)
Are LLM-based Evaluators Confusing NLG Quality Criteria?
von: Hu, Xinyu, et al.
Veröffentlicht: (2024)
von: Hu, Xinyu, et al.
Veröffentlicht: (2024)
Commentary on DHP concept
von: Joël Kruppa
Veröffentlicht: (2024)
von: Joël Kruppa
Veröffentlicht: (2024)
A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability
von: Hu, Xinyu, et al.
Veröffentlicht: (2025)
von: Hu, Xinyu, et al.
Veröffentlicht: (2025)
Analyzing and Evaluating Correlation Measures in NLG Meta-Evaluation
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
Large Language Models Are Active Critics in NLG Evaluation
von: Xu, Shuying, et al.
Veröffentlicht: (2024)
von: Xu, Shuying, et al.
Veröffentlicht: (2024)
OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs
von: Kartáč, Ivan, et al.
Veröffentlicht: (2025)
von: Kartáč, Ivan, et al.
Veröffentlicht: (2025)
LLM-based NLG Evaluation: Current Status and Challenges
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
Is Reference Necessary in the Evaluation of NLG Systems? When and Where?
von: Sheng, Shuqian, et al.
Veröffentlicht: (2024)
von: Sheng, Shuqian, et al.
Veröffentlicht: (2024)
NLG Evaluation: Past, Present, Future
von: Reiter, Ehud
Veröffentlicht: (2026)
von: Reiter, Ehud
Veröffentlicht: (2026)
Learning to Route LLMs with Confidence Tokens
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches
von: Yuan, Jiayi, et al.
Veröffentlicht: (2024)
von: Yuan, Jiayi, et al.
Veröffentlicht: (2024)
AmbigNLG: Addressing Task Ambiguity in Instruction for NLG
von: Niwa, Ayana, et al.
Veröffentlicht: (2024)
von: Niwa, Ayana, et al.
Veröffentlicht: (2024)
Confident or Seek Stronger: Exploring Uncertainty-Based On-device LLM Routing From Benchmarking to Generalization
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2025)
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2025)
DHP-Mapping: A Dense Panoptic Mapping System with Hierarchical World Representation and Label Optimization Techniques
von: Hu, Tianshuai, et al.
Veröffentlicht: (2024)
von: Hu, Tianshuai, et al.
Veröffentlicht: (2024)
LTSM-Bundle: A Toolbox and Benchmark on Large Language Models for Time Series Forecasting
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
Considerations for Cloud Security Operations
von: Cusick, James
Veröffentlicht: (2016)
von: Cusick, James
Veröffentlicht: (2016)
P‐2.3: Development Prospects and Current Status of Deep Learning Neural Network‐based Facial Capture in the Metaverse Field
von: Hongyu Qin, et al.
Veröffentlicht: (2024)
von: Hongyu Qin, et al.
Veröffentlicht: (2024)
DHP: Efficient Scaling of MLLM Training with Dynamic Hybrid Parallelism
von: Niu, Yifan, et al.
Veröffentlicht: (2026)
von: Niu, Yifan, et al.
Veröffentlicht: (2026)
DHP: Discrete Hierarchical Planning for Hierarchical Reinforcement Learning Agents
von: Sharma, Shashank, et al.
Veröffentlicht: (2025)
von: Sharma, Shashank, et al.
Veröffentlicht: (2025)
Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability
von: Hu, Xinyu, et al.
Veröffentlicht: (2024)
von: Hu, Xinyu, et al.
Veröffentlicht: (2024)
MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs
von: Zhang, Mengyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Mengyuan, et al.
Veröffentlicht: (2024)
Learning to Compress Prompt in Natural Language Formats
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
Defining and Detecting Vulnerability in Human Evaluation Guidelines: A Preliminary Study Towards Reliable NLG Evaluation
von: Ruan, Jie, et al.
Veröffentlicht: (2024)
von: Ruan, Jie, et al.
Veröffentlicht: (2024)
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
von: Jiang, Tianxiang, et al.
Veröffentlicht: (2025)
von: Jiang, Tianxiang, et al.
Veröffentlicht: (2025)
Taylor Unswift: Secured Weight Release for Large Language Models via Taylor Expansion
von: Wang, Guanchu, et al.
Veröffentlicht: (2024)
von: Wang, Guanchu, et al.
Veröffentlicht: (2024)
How to Select Datapoints for Efficient Human Evaluation of NLG Models?
von: Zouhar, Vilém, et al.
Veröffentlicht: (2025)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2025)
Leveraging Large Language Models for NLG Evaluation: Advances and Challenges
von: Li, Zhen, et al.
Veröffentlicht: (2024)
von: Li, Zhen, et al.
Veröffentlicht: (2024)
Not All Metrics Are Guilty: Improving NLG Evaluation by Diversifying References
von: Tang, Tianyi, et al.
Veröffentlicht: (2023)
von: Tang, Tianyi, et al.
Veröffentlicht: (2023)
Self-ensemble: Mitigating Confidence Mis-calibration for Large Language Models
von: Xu, Zicheng, et al.
Veröffentlicht: (2025)
von: Xu, Zicheng, et al.
Veröffentlicht: (2025)
TailNLG: A Multilingual Benchmark Addressing Verbalization of Long-Tail Entities
von: Draetta, Lia, et al.
Veröffentlicht: (2026)
von: Draetta, Lia, et al.
Veröffentlicht: (2026)
The Science of Evaluating Foundation Models
von: Yuan, Jiayi, et al.
Veröffentlicht: (2025)
von: Yuan, Jiayi, et al.
Veröffentlicht: (2025)
LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
von: Liusie, Adian, et al.
Veröffentlicht: (2023)
von: Liusie, Adian, et al.
Veröffentlicht: (2023)
Static Security Vulnerability Scanning of Proprietary and Open-Source Software: An Adaptable Process with Variants and Results
von: Cusick, James J.
Veröffentlicht: (2025)
von: Cusick, James J.
Veröffentlicht: (2025)
The First 50 Years of Software Reliability Engineering: A History of SRE with First Person Accounts
von: Cusick, James J.
Veröffentlicht: (2019)
von: Cusick, James J.
Veröffentlicht: (2019)
A Process To Support Cloud Release Preparation
von: Cusick, James J.
Veröffentlicht: (2022)
von: Cusick, James J.
Veröffentlicht: (2022)
An Object Web Seminar: A Retrospective on a Technical Dialogue Still Reverberating
von: Cusick, James J.
Veröffentlicht: (2026)
von: Cusick, James J.
Veröffentlicht: (2026)
Recursions for quadratic rotation symmetric functions weights
von: Cusick, Thomas W.
Veröffentlicht: (2025)
von: Cusick, Thomas W.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FairLENS: Assessing Fairness in Law Enforcement Speech Recognition
von: Wang, Yicheng, et al.
Veröffentlicht: (2024) -
Auto-Drafting Police Reports from Noisy ASR Outputs: A Trust-Centered LLM Approach
von: Kulkarni, Param, et al.
Veröffentlicht: (2025) -
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
von: Chang, Jiayi, et al.
Veröffentlicht: (2025) -
Are LLM-based Evaluators Confusing NLG Quality Criteria?
von: Hu, Xinyu, et al.
Veröffentlicht: (2024) -
Commentary on DHP concept
von: Joël Kruppa
Veröffentlicht: (2024)