Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mansurov, Jonibek, Sakip, Akhmed, Aji, Alham Fikri
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913873715527680
author Mansurov, Jonibek
Sakip, Akhmed
Aji, Alham Fikri
author_facet Mansurov, Jonibek
Sakip, Akhmed
Aji, Alham Fikri
contents In this paper, we show that knowledge distillation can be subverted to manipulate language model benchmark scores, revealing a critical vulnerability in current evaluation practices. We introduce "Data Laundering," a process that enables the covert transfer of benchmark-specific knowledge through seemingly legitimate intermediate training steps. Through extensive experiments with a 2-layer BERT student model, we show how this approach can achieve substantial improvements in benchmark accuracy (up to 75\% on GPQA) without developing genuine reasoning capabilities. Notably, this method can be exploited intentionally or even unintentionally, as researchers may inadvertently adopt this method and inflate scores without realising the implications. While our findings demonstrate the effectiveness of this technique, we present them as a cautionary tale highlighting the urgent need for more robust evaluation methods in AI. This work aims to contribute to the ongoing discussion about evaluation integrity in AI development and the need for benchmarks that more accurately reflect true model capabilities. The code is available at https://github.com/mbzuai-nlp/data_laundering.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15255
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation
Mansurov, Jonibek
Sakip, Akhmed
Aji, Alham Fikri
Computation and Language
Artificial Intelligence
In this paper, we show that knowledge distillation can be subverted to manipulate language model benchmark scores, revealing a critical vulnerability in current evaluation practices. We introduce "Data Laundering," a process that enables the covert transfer of benchmark-specific knowledge through seemingly legitimate intermediate training steps. Through extensive experiments with a 2-layer BERT student model, we show how this approach can achieve substantial improvements in benchmark accuracy (up to 75\% on GPQA) without developing genuine reasoning capabilities. Notably, this method can be exploited intentionally or even unintentionally, as researchers may inadvertently adopt this method and inflate scores without realising the implications. While our findings demonstrate the effectiveness of this technique, we present them as a cautionary tale highlighting the urgent need for more robust evaluation methods in AI. This work aims to contribute to the ongoing discussion about evaluation integrity in AI development and the need for benchmarks that more accurately reflect true model capabilities. The code is available at https://github.com/mbzuai-nlp/data_laundering.
title Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2412.15255