TreeEval: Benchmark-Free Evaluation of Large Language Models through Tree Planning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Xiang, Lan, Yunshi, Yang, Chao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915062295298048
author Li, Xiang
Lan, Yunshi
Yang, Chao
author_facet Li, Xiang
Lan, Yunshi
Yang, Chao
contents Recently, numerous new benchmarks have been established to evaluate the performance of large language models (LLMs) via either computing a holistic score or employing another LLM as a judge. However, these approaches suffer from data leakage due to the open access of the benchmark and inflexible evaluation process. To address this issue, we introduce $\textbf{TreeEval}$, a benchmark-free evaluation method for LLMs that let a high-performance LLM host an irreproducible evaluation session and essentially avoids the data leakage. Moreover, this LLM performs as an examiner to raise up a series of questions under a topic with a tree planing strategy, which considers the current evaluation status to decide the next question generation and ensures the completeness and efficiency of the evaluation process. We evaluate $6$ models of different parameter sizes, including $7$B, $13$B, and $33$B, and ultimately achieved the highest correlation coefficient with AlpacaEval2.0 using only around $45$ questions. We also conduct more analysis to show the robustness and reliability of TreeEval. Our code can be accessed via the provided https://github.com/Ashura5/TreeEval.
format Preprint
id arxiv_https___arxiv_org_abs_2402_13125
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TreeEval: Benchmark-Free Evaluation of Large Language Models through Tree Planning
Li, Xiang
Lan, Yunshi
Yang, Chao
Computation and Language
Artificial Intelligence
Recently, numerous new benchmarks have been established to evaluate the performance of large language models (LLMs) via either computing a holistic score or employing another LLM as a judge. However, these approaches suffer from data leakage due to the open access of the benchmark and inflexible evaluation process. To address this issue, we introduce $\textbf{TreeEval}$, a benchmark-free evaluation method for LLMs that let a high-performance LLM host an irreproducible evaluation session and essentially avoids the data leakage. Moreover, this LLM performs as an examiner to raise up a series of questions under a topic with a tree planing strategy, which considers the current evaluation status to decide the next question generation and ensures the completeness and efficiency of the evaluation process. We evaluate $6$ models of different parameter sizes, including $7$B, $13$B, and $33$B, and ultimately achieved the highest correlation coefficient with AlpacaEval2.0 using only around $45$ questions. We also conduct more analysis to show the robustness and reliability of TreeEval. Our code can be accessed via the provided https://github.com/Ashura5/TreeEval.
title TreeEval: Benchmark-Free Evaluation of Large Language Models through Tree Planning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2402.13125