DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Kaijie, Chen, Jiaao, Wang, Jindong, Gong, Neil Zhenqiang, Yang, Diyi, Xie, Xing
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909135953461248
author Zhu, Kaijie
Chen, Jiaao
Wang, Jindong
Gong, Neil Zhenqiang
Yang, Diyi
Xie, Xing
author_facet Zhu, Kaijie
Chen, Jiaao
Wang, Jindong
Gong, Neil Zhenqiang
Yang, Diyi
Xie, Xing
contents Large language models (LLMs) have achieved remarkable performance in various evaluation benchmarks. However, concerns are raised about potential data contamination in their considerable volume of training corpus. Moreover, the static nature and fixed complexity of current benchmarks may inadequately gauge the advancing capabilities of LLMs. In this paper, we introduce DyVal, a general and flexible protocol for dynamic evaluation of LLMs. Based on our framework, we build graph-informed DyVal by leveraging the structural advantage of directed acyclic graphs to dynamically generate evaluation samples with controllable complexities. DyVal generates challenging evaluation sets on reasoning tasks including mathematics, logical reasoning, and algorithm problems. We evaluate various LLMs ranging from Flan-T5-large to GPT-3.5-Turbo and GPT-4. Experiments show that LLMs perform worse in DyVal-generated evaluation samples with different complexities, highlighting the significance of dynamic evaluation. We also analyze the failure cases and results of different prompting methods. Moreover, DyVal-generated samples are not only evaluation sets, but also helpful data for fine-tuning to improve the performance of LLMs on existing benchmarks. We hope that DyVal can shed light on future evaluation research of LLMs. Code is available at: https://github.com/microsoft/promptbench.
format Preprint
id arxiv_https___arxiv_org_abs_2309_17167
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks
Zhu, Kaijie
Chen, Jiaao
Wang, Jindong
Gong, Neil Zhenqiang
Yang, Diyi
Xie, Xing
Artificial Intelligence
Computation and Language
Machine Learning
Large language models (LLMs) have achieved remarkable performance in various evaluation benchmarks. However, concerns are raised about potential data contamination in their considerable volume of training corpus. Moreover, the static nature and fixed complexity of current benchmarks may inadequately gauge the advancing capabilities of LLMs. In this paper, we introduce DyVal, a general and flexible protocol for dynamic evaluation of LLMs. Based on our framework, we build graph-informed DyVal by leveraging the structural advantage of directed acyclic graphs to dynamically generate evaluation samples with controllable complexities. DyVal generates challenging evaluation sets on reasoning tasks including mathematics, logical reasoning, and algorithm problems. We evaluate various LLMs ranging from Flan-T5-large to GPT-3.5-Turbo and GPT-4. Experiments show that LLMs perform worse in DyVal-generated evaluation samples with different complexities, highlighting the significance of dynamic evaluation. We also analyze the failure cases and results of different prompting methods. Moreover, DyVal-generated samples are not only evaluation sets, but also helpful data for fine-tuning to improve the performance of LLMs on existing benchmarks. We hope that DyVal can shed light on future evaluation research of LLMs. Code is available at: https://github.com/microsoft/promptbench.
title DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2309.17167