Benchmarking Distilled Language Models: Performance and Efficiency in Resource-Constrained Settings

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wani, Sachin Gopal, Page, Eric, Dholakia, Ajay, Ellison, David
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912921245712384
author Wani, Sachin Gopal
Page, Eric
Dholakia, Ajay
Ellison, David
author_facet Wani, Sachin Gopal
Page, Eric
Dholakia, Ajay
Ellison, David
contents Knowledge distillation offers a transformative pathway to developing powerful, yet efficient, small language models (SLMs) suitable for resource-constrained environments. In this paper, we benchmark the performance and computational cost of distilled models against their vanilla and proprietary counterparts, providing a quantitative analysis of their efficiency. Our results demonstrate that distillation creates a superior performance-tocompute curve. We find that creating a distilled 8B model is over 2,000 times more compute-efficient than training its vanilla counterpart, while achieving reasoning capabilities on par with, or even exceeding, standard models ten times its size. These findings validate distillation not just as a compression technique, but as a primary strategy for building state-of-the-art, accessible AI
format Preprint
id arxiv_https___arxiv_org_abs_2602_20164
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Benchmarking Distilled Language Models: Performance and Efficiency in Resource-Constrained Settings
Wani, Sachin Gopal
Page, Eric
Dholakia, Ajay
Ellison, David
Computation and Language
Machine Learning
Knowledge distillation offers a transformative pathway to developing powerful, yet efficient, small language models (SLMs) suitable for resource-constrained environments. In this paper, we benchmark the performance and computational cost of distilled models against their vanilla and proprietary counterparts, providing a quantitative analysis of their efficiency. Our results demonstrate that distillation creates a superior performance-tocompute curve. We find that creating a distilled 8B model is over 2,000 times more compute-efficient than training its vanilla counterpart, while achieving reasoning capabilities on par with, or even exceeding, standard models ten times its size. These findings validate distillation not just as a compression technique, but as a primary strategy for building state-of-the-art, accessible AI
title Benchmarking Distilled Language Models: Performance and Efficiency in Resource-Constrained Settings
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2602.20164