S3Eval: A Synthetic, Scalable, Systematic Evaluation Suite for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lei, Fangyu, Liu, Qian, Huang, Yiming, He, Shizhu, Zhao, Jun, Liu, Kang
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910400932478976
author Lei, Fangyu
Liu, Qian
Huang, Yiming
He, Shizhu
Zhao, Jun
Liu, Kang
author_facet Lei, Fangyu
Liu, Qian
Huang, Yiming
He, Shizhu
Zhao, Jun
Liu, Kang
contents The rapid development of Large Language Models (LLMs) has led to great strides in model capabilities like long-context understanding and reasoning. However, as LLMs are able to process longer contexts, it becomes more challenging to evaluate whether they have acquired certain capabilities, since the length of text (e.g., 200K tokens) they can process far exceeds what humans can reliably assess in a reasonable duration. In this paper, we propose using complex synthetic tasks as a proxy evaluation method, and present S3Eval, a Synthetic, Scalable, Systematic evaluation suite for LLMs evaluation. The synthetic nature of S3Eval provides users full control over the dataset, allowing them to systematically probe LLM capabilities by scaling text length and varying task difficulty across diverse scenarios. The strong correlation between S3Eval and real-world benchmarks demonstrates the soundness of using S3Eval for evaluation of LLMs. S3Eval provides a flexible and infinite long-context data generation method. We have generated a comprehensive dataset called S3Eval-Standard, and experimental results have shown that it poses significant challenges for all existing LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2310_15147
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle S3Eval: A Synthetic, Scalable, Systematic Evaluation Suite for Large Language Models
Lei, Fangyu
Liu, Qian
Huang, Yiming
He, Shizhu
Zhao, Jun
Liu, Kang
Computation and Language
The rapid development of Large Language Models (LLMs) has led to great strides in model capabilities like long-context understanding and reasoning. However, as LLMs are able to process longer contexts, it becomes more challenging to evaluate whether they have acquired certain capabilities, since the length of text (e.g., 200K tokens) they can process far exceeds what humans can reliably assess in a reasonable duration. In this paper, we propose using complex synthetic tasks as a proxy evaluation method, and present S3Eval, a Synthetic, Scalable, Systematic evaluation suite for LLMs evaluation. The synthetic nature of S3Eval provides users full control over the dataset, allowing them to systematically probe LLM capabilities by scaling text length and varying task difficulty across diverse scenarios. The strong correlation between S3Eval and real-world benchmarks demonstrates the soundness of using S3Eval for evaluation of LLMs. S3Eval provides a flexible and infinite long-context data generation method. We have generated a comprehensive dataset called S3Eval-Standard, and experimental results have shown that it poses significant challenges for all existing LLMs.
title S3Eval: A Synthetic, Scalable, Systematic Evaluation Suite for Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2310.15147