Revitalizing Saturated Benchmarks: A Weighted Metric Approach for Differentiating Large Language Model Performance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Etzine, Bryan, Hashemi, Masoud, Madhusudhan, Nishanth, Davasam, Sagar, Sharma, Roshnee, Madhusudhan, Sathwik Tejaswi, Yadav, Vikas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916646511181824
author Etzine, Bryan
Hashemi, Masoud
Madhusudhan, Nishanth
Davasam, Sagar
Sharma, Roshnee
Madhusudhan, Sathwik Tejaswi
Yadav, Vikas
author_facet Etzine, Bryan
Hashemi, Masoud
Madhusudhan, Nishanth
Davasam, Sagar
Sharma, Roshnee
Madhusudhan, Sathwik Tejaswi
Yadav, Vikas
contents Existing benchmarks are becoming saturated and struggle to separate model performances due to factors like data contamination and advancing LLM capabilities. This paper introduces EMDM (Enhanced Model Differentiation Metric), a novel weighted metric that revitalizes benchmarks by enhancing model separation. EMDM integrates final answer and Chain-of-Thought (CoT) reasoning correctness, assigning weights based on the complexity and reasoning depth required to solve a given sample in the evaluation data. Using a baseline LLM in two setups-Unguided, where the model has no prior exposure to test samples, and Guided, where the model has prior knowledge of the desired answer-EMDM distinguishes instances of varying difficulty. The CoT and answer correctness from these setups inform an optimization objective for weight assignment, resulting in a more nuanced evaluation of model performance. Compared to the exact match (EM) metric, which achieves 17% separation on ARC-Challenge, EMDM achieves 46%, demonstrating its effectiveness in differentiating models based on reasoning and knowledge requirements.
format Preprint
id arxiv_https___arxiv_org_abs_2503_05551
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revitalizing Saturated Benchmarks: A Weighted Metric Approach for Differentiating Large Language Model Performance
Etzine, Bryan
Hashemi, Masoud
Madhusudhan, Nishanth
Davasam, Sagar
Sharma, Roshnee
Madhusudhan, Sathwik Tejaswi
Yadav, Vikas
Machine Learning
Existing benchmarks are becoming saturated and struggle to separate model performances due to factors like data contamination and advancing LLM capabilities. This paper introduces EMDM (Enhanced Model Differentiation Metric), a novel weighted metric that revitalizes benchmarks by enhancing model separation. EMDM integrates final answer and Chain-of-Thought (CoT) reasoning correctness, assigning weights based on the complexity and reasoning depth required to solve a given sample in the evaluation data. Using a baseline LLM in two setups-Unguided, where the model has no prior exposure to test samples, and Guided, where the model has prior knowledge of the desired answer-EMDM distinguishes instances of varying difficulty. The CoT and answer correctness from these setups inform an optimization objective for weight assignment, resulting in a more nuanced evaluation of model performance. Compared to the exact match (EM) metric, which achieves 17% separation on ARC-Challenge, EMDM achieves 46%, demonstrating its effectiveness in differentiating models based on reasoning and knowledge requirements.
title Revitalizing Saturated Benchmarks: A Weighted Metric Approach for Differentiating Large Language Model Performance
topic Machine Learning
url https://arxiv.org/abs/2503.05551