Saved in:
| Main Authors: | Kovatchev, Venelin, Lease, Matthew |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2404.00748 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Finding Pareto Trade-offs in Fair and Accurate Detection of Toxic Speech
by: Gupta, Soumyajit, et al.
Published: (2022)
by: Gupta, Soumyajit, et al.
Published: (2022)
Capturing Classic Authorial Style in Long-Form Story Generation with GRPO Fine-Tuning
by: Liu, Jinlong, et al.
Published: (2025)
by: Liu, Jinlong, et al.
Published: (2025)
Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks
by: Luo, Zhimeng, et al.
Published: (2025)
by: Luo, Zhimeng, et al.
Published: (2025)
Transparent Screening for LLM Inference and Training Impacts
by: Pachot, Arnault, et al.
Published: (2026)
by: Pachot, Arnault, et al.
Published: (2026)
Codenames as a Benchmark for Large Language Models
by: Stephenson, Matthew, et al.
Published: (2024)
by: Stephenson, Matthew, et al.
Published: (2024)
Transparent Reference-free Automated Evaluation of Open-Ended User Survey Responses
by: An, Subin, et al.
Published: (2025)
by: An, Subin, et al.
Published: (2025)
Wrapper Boxes: Faithful Attribution of Model Predictions to Training Data
by: Su, Yiheng, et al.
Published: (2023)
by: Su, Yiheng, et al.
Published: (2023)
WXImpactBench: A Disruptive Weather Impact Understanding Benchmark for Evaluating Large Language Models
by: Yu, Yongan, et al.
Published: (2025)
by: Yu, Yongan, et al.
Published: (2025)
HoH: A Dynamic Benchmark for Evaluating the Impact of Outdated Information on Retrieval-Augmented Generation
by: Ouyang, Jie, et al.
Published: (2025)
by: Ouyang, Jie, et al.
Published: (2025)
Measuring the Impact of Lexical Training Data Coverage on Hallucination Detection in Large Language Models
by: Zhang, Shuo, et al.
Published: (2025)
by: Zhang, Shuo, et al.
Published: (2025)
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
by: Testini, Irene, et al.
Published: (2025)
by: Testini, Irene, et al.
Published: (2025)
Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions
by: Murugadoss, Bhuvanashree, et al.
Published: (2024)
by: Murugadoss, Bhuvanashree, et al.
Published: (2024)
Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic
by: Ingimundarson, Finnur Ágúst, et al.
Published: (2026)
by: Ingimundarson, Finnur Ágúst, et al.
Published: (2026)
Benchmarking Data Science Agents
by: Zhang, Yuge, et al.
Published: (2024)
by: Zhang, Yuge, et al.
Published: (2024)
Transparent and Coherent Procedural Mistake Detection
by: Storks, Shane, et al.
Published: (2024)
by: Storks, Shane, et al.
Published: (2024)
NarraBench: A Comprehensive Framework for Narrative Benchmarking
by: Hamilton, Sil, et al.
Published: (2025)
by: Hamilton, Sil, et al.
Published: (2025)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
by: Jiang, Botian, et al.
Published: (2024)
by: Jiang, Botian, et al.
Published: (2024)
DHP Benchmark: Are LLMs Good NLG Evaluators?
by: Wang, Yicheng, et al.
Published: (2024)
by: Wang, Yicheng, et al.
Published: (2024)
Generating Benchmarks for Factuality Evaluation of Language Models
by: Muhlgay, Dor, et al.
Published: (2023)
by: Muhlgay, Dor, et al.
Published: (2023)
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
by: Lunardi, Riccardo, et al.
Published: (2025)
by: Lunardi, Riccardo, et al.
Published: (2025)
The Qiyas Benchmark: Measuring ChatGPT Mathematical and Language Understanding in Arabic
by: Al-Khalifa, Shahad, et al.
Published: (2024)
by: Al-Khalifa, Shahad, et al.
Published: (2024)
Measuring and Benchmarking Large Language Models' Capabilities to Generate Persuasive Language
by: Pauli, Amalie Brogaard, et al.
Published: (2024)
by: Pauli, Amalie Brogaard, et al.
Published: (2024)
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
by: Bean, Andrew M., et al.
Published: (2025)
by: Bean, Andrew M., et al.
Published: (2025)
RUVA: Personalized Transparent On-Device Graph Reasoning
by: Conte, Gabriele, et al.
Published: (2026)
by: Conte, Gabriele, et al.
Published: (2026)
Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2025)
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2025)
The Self-Execution Benchmark: Measuring LLMs' Attempts to Overcome Their Lack of Self-Execution
by: Ezra, Elon, et al.
Published: (2025)
by: Ezra, Elon, et al.
Published: (2025)
QGEval: Benchmarking Multi-dimensional Evaluation for Question Generation
by: Fu, Weiping, et al.
Published: (2024)
by: Fu, Weiping, et al.
Published: (2024)
CEval: A Benchmark for Evaluating Counterfactual Text Generation
by: Nguyen, Van Bach, et al.
Published: (2024)
by: Nguyen, Van Bach, et al.
Published: (2024)
Generating Leakage-Free Benchmarks for Robust RAG Evaluation
by: Liu, Jiayi, et al.
Published: (2026)
by: Liu, Jiayi, et al.
Published: (2026)
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
by: Moore, Robert J., et al.
Published: (2026)
by: Moore, Robert J., et al.
Published: (2026)
RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation
by: Xu, Xinnuo, et al.
Published: (2025)
by: Xu, Xinnuo, et al.
Published: (2025)
Evaluating the Performance of Large Language Models on GAOKAO Benchmark
by: Zhang, Xiaotian, et al.
Published: (2023)
by: Zhang, Xiaotian, et al.
Published: (2023)
Kinship Data Benchmark for Multi-hop Reasoning
by: Sun, Tianda, et al.
Published: (2026)
by: Sun, Tianda, et al.
Published: (2026)
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
by: Lei, Fangyu, et al.
Published: (2025)
by: Lei, Fangyu, et al.
Published: (2025)
SPARQL Query Generation with LLMs: Measuring the Impact of Training Data Memorization and Knowledge Injection
by: Gashkov, Aleksandr, et al.
Published: (2025)
by: Gashkov, Aleksandr, et al.
Published: (2025)
An LLM Maturity Model for Reliable and Transparent Text-to-Query
by: Yu, Lei, et al.
Published: (2024)
by: Yu, Lei, et al.
Published: (2024)
Every Answer Matters: Evaluating Commonsense with Probabilistic Measures
by: Cheng, Qi, et al.
Published: (2024)
by: Cheng, Qi, et al.
Published: (2024)
AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents
by: Gioacchini, Luca, et al.
Published: (2024)
by: Gioacchini, Luca, et al.
Published: (2024)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
by: Deng, Shihan, et al.
Published: (2024)
by: Deng, Shihan, et al.
Published: (2024)
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
by: Yadav, Ankit, et al.
Published: (2024)
by: Yadav, Ankit, et al.
Published: (2024)
Similar Items
-
Finding Pareto Trade-offs in Fair and Accurate Detection of Toxic Speech
by: Gupta, Soumyajit, et al.
Published: (2022) -
Capturing Classic Authorial Style in Long-Form Story Generation with GRPO Fine-Tuning
by: Liu, Jinlong, et al.
Published: (2025) -
Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks
by: Luo, Zhimeng, et al.
Published: (2025) -
Transparent Screening for LLM Inference and Training Impacts
by: Pachot, Arnault, et al.
Published: (2026) -
Codenames as a Benchmark for Large Language Models
by: Stephenson, Matthew, et al.
Published: (2024)