Saved in:
Bibliographic Details
Main Author: Bradley, William F.
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2411.01533
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913843333038080
author Bradley, William F.
author_facet Bradley, William F.
contents As large language models (LLMs) become increasingly powerful, traditional evaluation metrics tend to saturate, making it challenging to distinguish between models. We propose a general method to transform existing LLM evaluations into a series of progressively more difficult tasks. These enhanced evaluations emphasize reasoning capabilities and can reveal relative performance differences that are not apparent in the original assessments. To demonstrate the effectiveness of our approach, we create a new multiple-choice test corpus, extend it into a family of evaluations, and assess a collection of LLMs. Our results offer insights into the comparative abilities of these models, particularly highlighting the differences between base LLMs and more recent "reasoning" models.
format Preprint
id arxiv_https___arxiv_org_abs_2411_01533
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing LLM Evaluations: The Garbling Trick
Bradley, William F.
Computation and Language
Artificial Intelligence
Machine Learning
As large language models (LLMs) become increasingly powerful, traditional evaluation metrics tend to saturate, making it challenging to distinguish between models. We propose a general method to transform existing LLM evaluations into a series of progressively more difficult tasks. These enhanced evaluations emphasize reasoning capabilities and can reveal relative performance differences that are not apparent in the original assessments. To demonstrate the effectiveness of our approach, we create a new multiple-choice test corpus, extend it into a family of evaluations, and assess a collection of LLMs. Our results offer insights into the comparative abilities of these models, particularly highlighting the differences between base LLMs and more recent "reasoning" models.
title Enhancing LLM Evaluations: The Garbling Trick
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.01533