Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Han, Yi, Xiaoyuan, Wei, Zhihua, Xiao, Ziang, Wang, Shu, Xie, Xing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912423062011904
author Jiang, Han
Yi, Xiaoyuan
Wei, Zhihua
Xiao, Ziang
Wang, Shu
Xie, Xing
author_facet Jiang, Han
Yi, Xiaoyuan
Wei, Zhihua
Xiao, Ziang
Wang, Shu
Xie, Xing
contents Warning: Contains harmful model outputs. Despite significant advancements, the propensity of Large Language Models (LLMs) to generate harmful and unethical content poses critical challenges. Measuring value alignment of LLMs becomes crucial for their regulation and responsible deployment. Although numerous benchmarks have been constructed to assess social bias, toxicity, and ethical issues in LLMs, those static benchmarks suffer from evaluation chronoeffect, in which, as models rapidly evolve, existing benchmarks may leak into training data or become saturated, overestimating ever-developing LLMs. To tackle this problem, we propose GETA, a novel generative evolving testing approach based on adaptive testing methods in measurement theory. Unlike traditional adaptive testing methods that rely on a static test item pool, GETA probes the underlying moral boundaries of LLMs by dynamically generating test items tailored to model capability. GETA co-evolves with LLMs by learning a joint distribution of item difficulty and model value conformity, thus effectively addressing evaluation chronoeffect. We evaluated various popular LLMs with GETA and demonstrated that 1) GETA can dynamically create difficulty-tailored test items and 2) GETA's evaluation results are more consistent with models' performance on unseen OOD and i.i.d. items, laying the groundwork for future evaluation paradigms.
format Preprint
id arxiv_https___arxiv_org_abs_2406_14230
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing
Jiang, Han
Yi, Xiaoyuan
Wei, Zhihua
Xiao, Ziang
Wang, Shu
Xie, Xing
Computation and Language
Artificial Intelligence
Computers and Society
Warning: Contains harmful model outputs. Despite significant advancements, the propensity of Large Language Models (LLMs) to generate harmful and unethical content poses critical challenges. Measuring value alignment of LLMs becomes crucial for their regulation and responsible deployment. Although numerous benchmarks have been constructed to assess social bias, toxicity, and ethical issues in LLMs, those static benchmarks suffer from evaluation chronoeffect, in which, as models rapidly evolve, existing benchmarks may leak into training data or become saturated, overestimating ever-developing LLMs. To tackle this problem, we propose GETA, a novel generative evolving testing approach based on adaptive testing methods in measurement theory. Unlike traditional adaptive testing methods that rely on a static test item pool, GETA probes the underlying moral boundaries of LLMs by dynamically generating test items tailored to model capability. GETA co-evolves with LLMs by learning a joint distribution of item difficulty and model value conformity, thus effectively addressing evaluation chronoeffect. We evaluated various popular LLMs with GETA and demonstrated that 1) GETA can dynamically create difficulty-tailored test items and 2) GETA's evaluation results are more consistent with models' performance on unseen OOD and i.i.d. items, laying the groundwork for future evaluation paradigms.
title Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing
topic Computation and Language
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2406.14230