Benchmarking the Energy Savings with Speculative Decoding Strategies
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866908823505076224 |
|---|---|
| author | Dutta, Rohit Koley, Paramita Poddar, Soham Misra, Janardan Podder, Sanjay Balani, Naveen Ghosh, Saptarshi Ganguly, Niloy |
| author_facet | Dutta, Rohit Koley, Paramita Poddar, Soham Misra, Janardan Podder, Sanjay Balani, Naveen Ghosh, Saptarshi Ganguly, Niloy |
| contents | Speculative decoding has emerged as an effective method to reduce latency and inference cost of LLM inferences. However, there has been inadequate attention towards the energy requirements of these models. To address this gap, this paper presents a comprehensive survey of energy requirements of speculative decoding strategies, with detailed analysis on how various factors -- model size and family, speculative decoding strategies, and dataset characteristics -- influence the energy optimizations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_09113 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Benchmarking the Energy Savings with Speculative Decoding Strategies Dutta, Rohit Koley, Paramita Poddar, Soham Misra, Janardan Podder, Sanjay Balani, Naveen Ghosh, Saptarshi Ganguly, Niloy Machine Learning Artificial Intelligence Computation and Language Speculative decoding has emerged as an effective method to reduce latency and inference cost of LLM inferences. However, there has been inadequate attention towards the energy requirements of these models. To address this gap, this paper presents a comprehensive survey of energy requirements of speculative decoding strategies, with detailed analysis on how various factors -- model size and family, speculative decoding strategies, and dataset characteristics -- influence the energy optimizations. |
| title | Benchmarking the Energy Savings with Speculative Decoding Strategies |
| topic | Machine Learning Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2602.09113 |