SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866916932958027776 |
|---|---|
| author | Pham, Nghiem Thanh Kieu, Tung Nguyen, Duc-Manh Xuan, Son Ha Duong-Trung, Nghia Le-Phuoc, Danh |
| author_facet | Pham, Nghiem Thanh Kieu, Tung Nguyen, Duc-Manh Xuan, Son Ha Duong-Trung, Nghia Le-Phuoc, Danh |
| contents | Small Language Models (SLMs) offer computational efficiency and accessibility, yet a systematic evaluation of their performance and environmental impact remains lacking. We introduce SLM-Bench, the first benchmark specifically designed to assess SLMs across multiple dimensions, including accuracy, computational efficiency, and sustainability metrics. SLM-Bench evaluates 15 SLMs on 9 NLP tasks using 23 datasets spanning 14 domains. The evaluation is conducted on 4 hardware configurations, providing a rigorous comparison of their effectiveness. Unlike prior benchmarks, SLM-Bench quantifies 11 metrics across correctness, computation, and consumption, enabling a holistic assessment of efficiency trade-offs. Our evaluation considers controlled hardware conditions, ensuring fair comparisons across models. We develop an open-source benchmarking pipeline with standardized evaluation protocols to facilitate reproducibility and further research. Our findings highlight the diverse trade-offs among SLMs, where some models excel in accuracy while others achieve superior energy efficiency. SLM-Bench sets a new standard for SLM evaluation, bridging the gap between resource efficiency and real-world applicability. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_15478 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version Pham, Nghiem Thanh Kieu, Tung Nguyen, Duc-Manh Xuan, Son Ha Duong-Trung, Nghia Le-Phuoc, Danh Computation and Language Computers and Society Performance Small Language Models (SLMs) offer computational efficiency and accessibility, yet a systematic evaluation of their performance and environmental impact remains lacking. We introduce SLM-Bench, the first benchmark specifically designed to assess SLMs across multiple dimensions, including accuracy, computational efficiency, and sustainability metrics. SLM-Bench evaluates 15 SLMs on 9 NLP tasks using 23 datasets spanning 14 domains. The evaluation is conducted on 4 hardware configurations, providing a rigorous comparison of their effectiveness. Unlike prior benchmarks, SLM-Bench quantifies 11 metrics across correctness, computation, and consumption, enabling a holistic assessment of efficiency trade-offs. Our evaluation considers controlled hardware conditions, ensuring fair comparisons across models. We develop an open-source benchmarking pipeline with standardized evaluation protocols to facilitate reproducibility and further research. Our findings highlight the diverse trade-offs among SLMs, where some models excel in accuracy while others achieve superior energy efficiency. SLM-Bench sets a new standard for SLM evaluation, bridging the gap between resource efficiency and real-world applicability. |
| title | SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version |
| topic | Computation and Language Computers and Society Performance |
| url | https://arxiv.org/abs/2508.15478 |