Evaluating the Utility of Grounding Documents with Reference-Free LLM-based Metrics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hua, Yilun, Castellucci, Giuseppe, Schulam, Peter, Elfardy, Heba, Small, Kevin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unveiling LLM Evaluation Focused on Metrics: Challenges and Solutions
von: Hu, Taojun, et al.
Veröffentlicht: (2024)
von: Hu, Taojun, et al.
Veröffentlicht: (2024)
ConCISE: A Reference-Free Conciseness Evaluation Metric for LLM-Generated Answers
von: Ghafari, Seyed Mohssen, et al.
Veröffentlicht: (2025)
von: Ghafari, Seyed Mohssen, et al.
Veröffentlicht: (2025)
TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness
von: Zheng, Danna, et al.
Veröffentlicht: (2024)
von: Zheng, Danna, et al.
Veröffentlicht: (2024)
Mitigating the Impact of Reference Quality on Evaluation of Summarization Systems with Reference-Free Metrics
von: Gigant, Théo, et al.
Veröffentlicht: (2024)
von: Gigant, Théo, et al.
Veröffentlicht: (2024)
Cobra Effect in Reference-Free Image Captioning Metrics
von: Ma, Zheng, et al.
Veröffentlicht: (2024)
von: Ma, Zheng, et al.
Veröffentlicht: (2024)
An Examination of the Robustness of Reference-Free Image Captioning Evaluation Metrics
von: Ahmadi, Saba, et al.
Veröffentlicht: (2023)
von: Ahmadi, Saba, et al.
Veröffentlicht: (2023)
SCORE: Specificity, Context Utilization, Robustness, and Relevance for Reference-Free LLM Evaluation
von: Shomee, Homaira Huda, et al.
Veröffentlicht: (2026)
von: Shomee, Homaira Huda, et al.
Veröffentlicht: (2026)
Reference-Free Evaluation of Taxonomies
von: Wullschleger, Pascal, et al.
Veröffentlicht: (2025)
von: Wullschleger, Pascal, et al.
Veröffentlicht: (2025)
Leveraging Interesting Facts to Enhance User Engagement with Conversational Interfaces
von: Vedula, Nikhita, et al.
Veröffentlicht: (2024)
von: Vedula, Nikhita, et al.
Veröffentlicht: (2024)
Rethinking Atomic Decomposition for LLM Judges: A Prompt-Controlled Study of Reference-Grounded QA Evaluation
von: Zhang, Xinran
Veröffentlicht: (2026)
von: Zhang, Xinran
Veröffentlicht: (2026)
Model Utility Law: Evaluating LLMs beyond Performance through Mechanism Interpretable Metric
von: Cao, Yixin, et al.
Veröffentlicht: (2025)
von: Cao, Yixin, et al.
Veröffentlicht: (2025)
Reference-free Evaluation Metrics for Text Generation: A Survey
von: Ito, Takumi, et al.
Veröffentlicht: (2025)
von: Ito, Takumi, et al.
Veröffentlicht: (2025)
Not All Metrics Are Guilty: Improving NLG Evaluation by Diversifying References
von: Tang, Tianyi, et al.
Veröffentlicht: (2023)
von: Tang, Tianyi, et al.
Veröffentlicht: (2023)
Talk Less, Interact Better: Evaluating In-context Conversational Adaptation in Multimodal LLMs
von: Hua, Yilun, et al.
Veröffentlicht: (2024)
von: Hua, Yilun, et al.
Veröffentlicht: (2024)
CRIMSON: A Clinically-Grounded LLM-Based Metric for Generative Radiology Report Evaluation
von: Baharoon, Mohammed, et al.
Veröffentlicht: (2026)
von: Baharoon, Mohammed, et al.
Veröffentlicht: (2026)
Readability Reconsidered: A Cross-Dataset Analysis of Reference-Free Metrics
von: Belem, Catarina G, et al.
Veröffentlicht: (2025)
von: Belem, Catarina G, et al.
Veröffentlicht: (2025)
FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model
von: Lee, Yebin, et al.
Veröffentlicht: (2024)
von: Lee, Yebin, et al.
Veröffentlicht: (2024)
KIEval: Evaluation Metric for Document Key Information Extraction
von: Khang, Minsoo, et al.
Veröffentlicht: (2025)
von: Khang, Minsoo, et al.
Veröffentlicht: (2025)
Measuring the Robustness of Reference-Free Dialogue Evaluation Systems
von: Vasselli, Justin, et al.
Veröffentlicht: (2025)
von: Vasselli, Justin, et al.
Veröffentlicht: (2025)
MILE-RefHumEval: A Reference-Free, Multi-Independent LLM Framework for Human-Aligned Evaluation
von: Srun, Nalin, et al.
Veröffentlicht: (2026)
von: Srun, Nalin, et al.
Veröffentlicht: (2026)
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2026)
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2026)
Rethinking Scientific Summarization Evaluation: Grounding Explainable Metrics on Facet-aware Benchmark
von: Chen, Xiuying, et al.
Veröffentlicht: (2024)
von: Chen, Xiuying, et al.
Veröffentlicht: (2024)
Generative Explore-Exploit: Training-free Optimization of Generative Recommender Systems using LLM Optimizers
von: Senel, Lütfi Kerem, et al.
Veröffentlicht: (2024)
von: Senel, Lütfi Kerem, et al.
Veröffentlicht: (2024)
Survey on Evaluation of LLM-based Agents
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025)
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
PromptOptMe: Error-Aware Prompt Compression for LLM-based MT Evaluation Metrics
von: Larionov, Daniil, et al.
Veröffentlicht: (2024)
von: Larionov, Daniil, et al.
Veröffentlicht: (2024)
CausalScore: An Automatic Reference-Free Metric for Assessing Response Relevance in Open-Domain Dialogue Systems
von: Feng, Tao, et al.
Veröffentlicht: (2024)
von: Feng, Tao, et al.
Veröffentlicht: (2024)
Beyond Single Ground Truth: Reference Monism as Epistemic Injustice in ASR Evaluation
von: Choi, Anna Seo Gyeong, et al.
Veröffentlicht: (2026)
von: Choi, Anna Seo Gyeong, et al.
Veröffentlicht: (2026)
LLM-Guided Planning and Summary-Based Scientific Text Simplification: DS@GT at CLEF 2025 SimpleText
von: Marturi, Krishna Chaitanya, et al.
Veröffentlicht: (2025)
von: Marturi, Krishna Chaitanya, et al.
Veröffentlicht: (2025)
DOCBENCH: A Benchmark for Evaluating LLM-based Document Reading Systems
von: Zou, Anni, et al.
Veröffentlicht: (2024)
von: Zou, Anni, et al.
Veröffentlicht: (2024)
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
von: Krumdick, Michael, et al.
Veröffentlicht: (2025)
von: Krumdick, Michael, et al.
Veröffentlicht: (2025)
Grounded in Reality: Learning and Deploying Proactive LLM from Offline Logs
von: Wei, Fei, et al.
Veröffentlicht: (2025)
von: Wei, Fei, et al.
Veröffentlicht: (2025)
Clinically Grounded Agent-based Report Evaluation: An Interpretable Metric for Radiology Report Generation
von: Dua, Radhika, et al.
Veröffentlicht: (2025)
von: Dua, Radhika, et al.
Veröffentlicht: (2025)
Enhancing Low-Resource LLMs Classification with PEFT and Synthetic Data
von: Patwa, Parth, et al.
Veröffentlicht: (2024)
von: Patwa, Parth, et al.
Veröffentlicht: (2024)
Planning Anything with Rigor: General-Purpose Zero-Shot Planning with LLM-based Formalized Programming
von: Hao, Yilun, et al.
Veröffentlicht: (2024)
von: Hao, Yilun, et al.
Veröffentlicht: (2024)
Theory-Grounded Evaluation Exposes the Authorship Gap in LLM Personalization
von: Sawant, Yash Ganpat
Veröffentlicht: (2026)
von: Sawant, Yash Ganpat
Veröffentlicht: (2026)
Reference-based Metrics Disprove Themselves in Question Generation
von: Nguyen, Bang, et al.
Veröffentlicht: (2024)
von: Nguyen, Bang, et al.
Veröffentlicht: (2024)
CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages
von: Yang, Yilun, et al.
Veröffentlicht: (2025)
von: Yang, Yilun, et al.
Veröffentlicht: (2025)
NovAScore: A New Automated Metric for Evaluating Document Level Novelty
von: Ai, Lin, et al.
Veröffentlicht: (2024)
von: Ai, Lin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Unveiling LLM Evaluation Focused on Metrics: Challenges and Solutions
von: Hu, Taojun, et al.
Veröffentlicht: (2024) -
ConCISE: A Reference-Free Conciseness Evaluation Metric for LLM-Generated Answers
von: Ghafari, Seyed Mohssen, et al.
Veröffentlicht: (2025) -
TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness
von: Zheng, Danna, et al.
Veröffentlicht: (2024) -
Mitigating the Impact of Reference Quality on Evaluation of Summarization Systems with Reference-Free Metrics
von: Gigant, Théo, et al.
Veröffentlicht: (2024) -
Cobra Effect in Reference-Free Image Captioning Metrics
von: Ma, Zheng, et al.
Veröffentlicht: (2024)