Honey, I Shrunk the Language Model: Impact of Knowledge Distillation Methods on Performance and Explainability

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hendriks, Daniel, Spitzer, Philipp, Kühl, Niklas, Satzger, Gerhard
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908332525092864
author Hendriks, Daniel
Spitzer, Philipp
Kühl, Niklas
Satzger, Gerhard
author_facet Hendriks, Daniel
Spitzer, Philipp
Kühl, Niklas
Satzger, Gerhard
contents Artificial Intelligence (AI) has increasingly influenced modern society, recently in particular through significant advancements in Large Language Models (LLMs). However, high computational and storage demands of LLMs still limit their deployment in resource-constrained environments. Knowledge distillation addresses this challenge by training a small student model from a larger teacher model. Previous research has introduced several distillation methods for both generating training data and for training the student model. Despite their relevance, the effects of state-of-the-art distillation methods on model performance and explainability have not been thoroughly investigated and compared. In this work, we enlarge the set of available methods by applying critique-revision prompting to distillation for data generation and by synthesizing existing methods for training. For these methods, we provide a systematic comparison based on the widely used Commonsense Question-Answering (CQA) dataset. While we measure performance via student model accuracy, we employ a human-grounded study to evaluate explainability. We contribute new distillation methods and their comparison in terms of both performance and explainability. This should further advance the distillation of small language models and, thus, contribute to broader applicability and faster diffusion of LLM technology.
format Preprint
id arxiv_https___arxiv_org_abs_2504_16056
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Honey, I Shrunk the Language Model: Impact of Knowledge Distillation Methods on Performance and Explainability
Hendriks, Daniel
Spitzer, Philipp
Kühl, Niklas
Satzger, Gerhard
Computation and Language
Artificial Intelligence (AI) has increasingly influenced modern society, recently in particular through significant advancements in Large Language Models (LLMs). However, high computational and storage demands of LLMs still limit their deployment in resource-constrained environments. Knowledge distillation addresses this challenge by training a small student model from a larger teacher model. Previous research has introduced several distillation methods for both generating training data and for training the student model. Despite their relevance, the effects of state-of-the-art distillation methods on model performance and explainability have not been thoroughly investigated and compared. In this work, we enlarge the set of available methods by applying critique-revision prompting to distillation for data generation and by synthesizing existing methods for training. For these methods, we provide a systematic comparison based on the widely used Commonsense Question-Answering (CQA) dataset. While we measure performance via student model accuracy, we employ a human-grounded study to evaluate explainability. We contribute new distillation methods and their comparison in terms of both performance and explainability. This should further advance the distillation of small language models and, thus, contribute to broader applicability and faster diffusion of LLM technology.
title Honey, I Shrunk the Language Model: Impact of Knowledge Distillation Methods on Performance and Explainability
topic Computation and Language
url https://arxiv.org/abs/2504.16056