Gespeichert in:
| Hauptverfasser: | Hu, Taojun, Zhou, Xiao-Hua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2404.09135 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM-based NLG Evaluation: Current Status and Challenges
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
Evaluating the Utility of Grounding Documents with Reference-Free LLM-based Metrics
von: Hua, Yilun, et al.
Veröffentlicht: (2026)
von: Hua, Yilun, et al.
Veröffentlicht: (2026)
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing
von: Long, Lingkun, et al.
Veröffentlicht: (2026)
von: Long, Lingkun, et al.
Veröffentlicht: (2026)
LLM-as-a-tutor in EFL Writing Education: Focusing on Evaluation of Student-LLM Interaction
von: Han, Jieun, et al.
Veröffentlicht: (2023)
von: Han, Jieun, et al.
Veröffentlicht: (2023)
Faithful Model Evaluation for Model-Based Metrics
von: Goyal, Palash, et al.
Veröffentlicht: (2023)
von: Goyal, Palash, et al.
Veröffentlicht: (2023)
Mind the Blind Spots: A Focus-Level Evaluation Framework for LLM Reviews
von: Shin, Hyungyu, et al.
Veröffentlicht: (2025)
von: Shin, Hyungyu, et al.
Veröffentlicht: (2025)
Unveiling Knowledge Utilization Mechanisms in LLM-based Retrieval-Augmented Generation
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
Automatic Evaluation Metrics for Document-level Translation: Overview, Challenges and Trends
von: GUO, Jiaxin, et al.
Veröffentlicht: (2025)
von: GUO, Jiaxin, et al.
Veröffentlicht: (2025)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
PsycoLLM: Enhancing LLM for Psychological Understanding and Evaluation
von: Hu, Jinpeng, et al.
Veröffentlicht: (2024)
von: Hu, Jinpeng, et al.
Veröffentlicht: (2024)
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
von: Wei, Hui, et al.
Veröffentlicht: (2024)
von: Wei, Hui, et al.
Veröffentlicht: (2024)
LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation
von: Eigler, Lukáš, et al.
Veröffentlicht: (2026)
von: Eigler, Lukáš, et al.
Veröffentlicht: (2026)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
Challenging the Evaluator: LLM Sycophancy Under User Rebuttal
von: Kim, Sungwon, et al.
Veröffentlicht: (2025)
von: Kim, Sungwon, et al.
Veröffentlicht: (2025)
A Study on Question-Answer Dataset for LLM Safety Evaluation with a Focus on Illegal Activities
von: Imamura, Kenji, et al.
Veröffentlicht: (2026)
von: Imamura, Kenji, et al.
Veröffentlicht: (2026)
GAMBIT+: A Challenge Set for Evaluating Gender Bias in Machine Translation Quality Estimation Metrics
von: Filandrianos, Giorgos, et al.
Veröffentlicht: (2025)
von: Filandrianos, Giorgos, et al.
Veröffentlicht: (2025)
A Survey of LLM-based Deep Search Agents: Paradigm, Optimization, Evaluation, and Challenges
von: Xi, Yunjia, et al.
Veröffentlicht: (2025)
von: Xi, Yunjia, et al.
Veröffentlicht: (2025)
How to Get Your LLM to Generate Challenging Problems for Evaluation
von: Patel, Arkil, et al.
Veröffentlicht: (2025)
von: Patel, Arkil, et al.
Veröffentlicht: (2025)
IDGen: Item Discrimination Induced Prompt Generation for LLM Evaluation
von: Lin, Fan, et al.
Veröffentlicht: (2024)
von: Lin, Fan, et al.
Veröffentlicht: (2024)
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
LLM Inference Unveiled: Survey and Roofline Model Insights
von: Yuan, Zhihang, et al.
Veröffentlicht: (2024)
von: Yuan, Zhihang, et al.
Veröffentlicht: (2024)
Contextual Metric Meta-Evaluation by Measuring Local Metric Accuracy
von: Deviyani, Athiya, et al.
Veröffentlicht: (2025)
von: Deviyani, Athiya, et al.
Veröffentlicht: (2025)
PromptOptMe: Error-Aware Prompt Compression for LLM-based MT Evaluation Metrics
von: Larionov, Daniil, et al.
Veröffentlicht: (2024)
von: Larionov, Daniil, et al.
Veröffentlicht: (2024)
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions
von: Tang, Bangsheng, et al.
Veröffentlicht: (2025)
von: Tang, Bangsheng, et al.
Veröffentlicht: (2025)
Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics
von: Cho, Yousang, et al.
Veröffentlicht: (2025)
von: Cho, Yousang, et al.
Veröffentlicht: (2025)
NLP for Local Governance Meeting Records: A Focus Article on Tasks, Datasets, Metrics and Benchmark
von: Campos, Ricardo, et al.
Veröffentlicht: (2026)
von: Campos, Ricardo, et al.
Veröffentlicht: (2026)
Developing a Multilingual Dataset and Evaluation Metrics for Code-Switching: A Focus on Hong Kong's Polylingual Dynamics
von: Xie, Peng, et al.
Veröffentlicht: (2023)
von: Xie, Peng, et al.
Veröffentlicht: (2023)
Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference
von: Luo, Zhifan, et al.
Veröffentlicht: (2025)
von: Luo, Zhifan, et al.
Veröffentlicht: (2025)
The Challenges of Evaluating LLM Applications: An Analysis of Automated, Human, and LLM-Based Approaches
von: Abeysinghe, Bhashithe, et al.
Veröffentlicht: (2024)
von: Abeysinghe, Bhashithe, et al.
Veröffentlicht: (2024)
Evaluating Compositional Approaches for Focus and Sentiment Analysis
von: Kellert, Olga, et al.
Veröffentlicht: (2025)
von: Kellert, Olga, et al.
Veröffentlicht: (2025)
ContrastScore: Towards Higher Quality, Less Biased, More Efficient Evaluation Metrics with Contrastive Evaluation
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Unveiling Language Competence Neurons: A Psycholinguistic Approach to Model Interpretability
von: Duan, Xufeng, et al.
Veröffentlicht: (2024)
von: Duan, Xufeng, et al.
Veröffentlicht: (2024)
Visualizing Uncertainty in Translation Tasks: An Evaluation of LLM Performance and Confidence Metrics
von: Park, Jin Hyun, et al.
Veröffentlicht: (2025)
von: Park, Jin Hyun, et al.
Veröffentlicht: (2025)
APPLS: Evaluating Evaluation Metrics for Plain Language Summarization
von: Guo, Yue, et al.
Veröffentlicht: (2023)
von: Guo, Yue, et al.
Veröffentlicht: (2023)
A Comparison of LLM Finetuning Methods & Evaluation Metrics with Travel Chatbot Use Case
von: Meyer, Sonia, et al.
Veröffentlicht: (2024)
von: Meyer, Sonia, et al.
Veröffentlicht: (2024)
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
von: Chang, Jiayi, et al.
Veröffentlicht: (2025)
von: Chang, Jiayi, et al.
Veröffentlicht: (2025)
Evaluating Metrics for Bias in Word Embeddings
von: Schröder, Sarah, et al.
Veröffentlicht: (2021)
von: Schröder, Sarah, et al.
Veröffentlicht: (2021)
ProverbEval: Exploring LLM Evaluation Challenges for Low-resource Language Understanding
von: Azime, Israel Abebe, et al.
Veröffentlicht: (2024)
von: Azime, Israel Abebe, et al.
Veröffentlicht: (2024)
A likelihood-based sensitivity analysis for addressing publication bias in meta-analysis of diagnostic studies using exact likelihood
von: Hu, Taojun, et al.
Veröffentlicht: (2024)
von: Hu, Taojun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LLM-based NLG Evaluation: Current Status and Challenges
von: Gao, Mingqi, et al.
Veröffentlicht: (2024) -
Evaluating the Utility of Grounding Documents with Reference-Free LLM-based Metrics
von: Hua, Yilun, et al.
Veröffentlicht: (2026) -
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025) -
Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing
von: Long, Lingkun, et al.
Veröffentlicht: (2026) -
LLM-as-a-tutor in EFL Writing Education: Focusing on Evaluation of Student-LLM Interaction
von: Han, Jieun, et al.
Veröffentlicht: (2023)