Benchmarking Distilled Language Models: Performance and Efficiency in Resource-Constrained Settings
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866912921245712384 |
|---|---|
| author | Wani, Sachin Gopal Page, Eric Dholakia, Ajay Ellison, David |
| author_facet | Wani, Sachin Gopal Page, Eric Dholakia, Ajay Ellison, David |
| contents | Knowledge distillation offers a transformative pathway to developing powerful, yet efficient, small language models (SLMs) suitable for resource-constrained environments. In this paper, we benchmark the performance and computational cost of distilled models against their vanilla and proprietary counterparts, providing a quantitative analysis of their efficiency. Our results demonstrate that distillation creates a superior performance-tocompute curve. We find that creating a distilled 8B model is over 2,000 times more compute-efficient than training its vanilla counterpart, while achieving reasoning capabilities on par with, or even exceeding, standard models ten times its size. These findings validate distillation not just as a compression technique, but as a primary strategy for building state-of-the-art, accessible AI |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_20164 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Benchmarking Distilled Language Models: Performance and Efficiency in Resource-Constrained Settings Wani, Sachin Gopal Page, Eric Dholakia, Ajay Ellison, David Computation and Language Machine Learning Knowledge distillation offers a transformative pathway to developing powerful, yet efficient, small language models (SLMs) suitable for resource-constrained environments. In this paper, we benchmark the performance and computational cost of distilled models against their vanilla and proprietary counterparts, providing a quantitative analysis of their efficiency. Our results demonstrate that distillation creates a superior performance-tocompute curve. We find that creating a distilled 8B model is over 2,000 times more compute-efficient than training its vanilla counterpart, while achieving reasoning capabilities on par with, or even exceeding, standard models ten times its size. These findings validate distillation not just as a compression technique, but as a primary strategy for building state-of-the-art, accessible AI |
| title | Benchmarking Distilled Language Models: Performance and Efficiency in Resource-Constrained Settings |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2602.20164 |