Benchmarking the Energy Savings with Speculative Decoding Strategies
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dutta, Rohit, Koley, Paramita, Poddar, Soham, Misra, Janardan, Podder, Sanjay, Balani, Naveen, Ghosh, Saptarshi, Ganguly, Niloy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Sustainable NLP: Insights from Benchmarking Inference Energy in Large Language Models
von: Poddar, Soham, et al.
Veröffentlicht: (2025)
von: Poddar, Soham, et al.
Veröffentlicht: (2025)
Brevity is the soul of sustainability: Characterizing LLM response lengths
von: Poddar, Soham, et al.
Veröffentlicht: (2025)
von: Poddar, Soham, et al.
Veröffentlicht: (2025)
Label-semantics Aware Generative Approach for Domain-Agnostic Multilabel Classification
von: Khatuya, Subhendu, et al.
Veröffentlicht: (2025)
von: Khatuya, Subhendu, et al.
Veröffentlicht: (2025)
RSTGCN: Railway-centric Spatio-Temporal Graph Convolutional Network for Train Delay Prediction
von: Chowdhury, Koyena, et al.
Veröffentlicht: (2025)
von: Chowdhury, Koyena, et al.
Veröffentlicht: (2025)
Advancing Decoding Strategies: Enhancements in Locally Typical Sampling for LLMs
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
Understanding Syllogistic Reasoning in LLMs from Formal and Natural Language Perspectives
von: Poddar, Aheli, et al.
Veröffentlicht: (2025)
von: Poddar, Aheli, et al.
Veröffentlicht: (2025)
Utilising Large Language Models for Generating Effective Counter Arguments to Anti-Vaccine Tweets
von: Dhanuka, Utsav, et al.
Veröffentlicht: (2025)
von: Dhanuka, Utsav, et al.
Veröffentlicht: (2025)
How COVID-19 has Impacted the Anti-Vaccine Discourse: A Large-Scale Twitter Study Spanning Pre-COVID and Post-COVID Era
von: Poddar, Soham, et al.
Veröffentlicht: (2024)
von: Poddar, Soham, et al.
Veröffentlicht: (2024)
Learning Faster with Better Tokens: Parameter-Efficient Vocabulary Adaptation for Specialized Text Summarization
von: Balde, Gunjan, et al.
Veröffentlicht: (2026)
von: Balde, Gunjan, et al.
Veröffentlicht: (2026)
The Disparate Impacts of Speculative Decoding
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
Constrained Decoding with Speculative Lookaheads
von: Nakshatri, Nishanth, et al.
Veröffentlicht: (2024)
von: Nakshatri, Nishanth, et al.
Veröffentlicht: (2024)
Scaling Laws for Speculative Decoding
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
Cross-Attention Speculative Decoding
von: Zhong, Wei, et al.
Veröffentlicht: (2025)
von: Zhong, Wei, et al.
Veröffentlicht: (2025)
Speculative Decoding: Performance or Illusion?
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025)
Mamba Drafters for Speculative Decoding
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
Goose: Anisotropic Speculation Trees for Training-Free Speculative Decoding
von: Jin, Tao, et al.
Veröffentlicht: (2026)
von: Jin, Tao, et al.
Veröffentlicht: (2026)
Dynamic Depth Decoding: Faster Speculative Decoding for LLMs
von: Brown, Oscar, et al.
Veröffentlicht: (2024)
von: Brown, Oscar, et al.
Veröffentlicht: (2024)
Applicability of Large Language Models and Generative Models for Legal Case Judgement Summarization
von: Deroy, Aniket, et al.
Veröffentlicht: (2024)
von: Deroy, Aniket, et al.
Veröffentlicht: (2024)
Batch Speculative Decoding Done Right
von: Zhang, Ranran Haoran, et al.
Veröffentlicht: (2025)
von: Zhang, Ranran Haoran, et al.
Veröffentlicht: (2025)
RASD: Retrieval-Augmented Speculative Decoding
von: Quan, Guofeng, et al.
Veröffentlicht: (2025)
von: Quan, Guofeng, et al.
Veröffentlicht: (2025)
Cacheback: Speculative Decoding With Nothing But Cache
von: Ma, Zhiyao, et al.
Veröffentlicht: (2025)
von: Ma, Zhiyao, et al.
Veröffentlicht: (2025)
Speculative Decoding for Multi-Sample Inference
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
Online Speculative Decoding
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
Hybrid Verified Decoding: Learning to Allocate Verification in Speculative Decoding
von: Su, Xin, et al.
Veröffentlicht: (2026)
von: Su, Xin, et al.
Veröffentlicht: (2026)
MILPaC: A Novel Benchmark for Evaluating Translation of Legal Text to Indian Languages
von: Mahapatra, Sayan, et al.
Veröffentlicht: (2023)
von: Mahapatra, Sayan, et al.
Veröffentlicht: (2023)
IDALC: A Semi-Supervised Framework for Intent Detection and Active Learning based Correction
von: Mullick, Ankan, et al.
Veröffentlicht: (2025)
von: Mullick, Ankan, et al.
Veröffentlicht: (2025)
ICPR 2024 Competition on Multilingual Claim-Span Identification
von: Poddar, Soham, et al.
Veröffentlicht: (2024)
von: Poddar, Soham, et al.
Veröffentlicht: (2024)
Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
von: Sun, Shengyin, et al.
Veröffentlicht: (2025)
von: Sun, Shengyin, et al.
Veröffentlicht: (2025)
Mixture of Attentions For Speculative Decoding
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
GRIFFIN: Effective Token Alignment for Faster Speculative Decoding
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
von: Zhang, Zihong, et al.
Veröffentlicht: (2026)
von: Zhang, Zihong, et al.
Veröffentlicht: (2026)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
Acceptance Dynamics Across Cognitive Domains in Speculative Decoding
von: Mahmoud, Saif
Veröffentlicht: (2026)
von: Mahmoud, Saif
Veröffentlicht: (2026)
Traversal Verification for Speculative Tree Decoding
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
Faster Cascades via Speculative Decoding
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025)
DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
When, What, and How: Rethinking Retrieval-Enhanced Speculative Decoding
von: Fang, Min, et al.
Veröffentlicht: (2025)
von: Fang, Min, et al.
Veröffentlicht: (2025)
PACER: Blockwise Pre-verification for Speculative Decoding with Adaptive Length
von: Zhang, Situo, et al.
Veröffentlicht: (2026)
von: Zhang, Situo, et al.
Veröffentlicht: (2026)
Entropy-Aware Speculative Decoding Toward Improved LLM Reasoning
von: Su, Tiancheng, et al.
Veröffentlicht: (2025)
von: Su, Tiancheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards Sustainable NLP: Insights from Benchmarking Inference Energy in Large Language Models
von: Poddar, Soham, et al.
Veröffentlicht: (2025) -
Brevity is the soul of sustainability: Characterizing LLM response lengths
von: Poddar, Soham, et al.
Veröffentlicht: (2025) -
Label-semantics Aware Generative Approach for Domain-Agnostic Multilabel Classification
von: Khatuya, Subhendu, et al.
Veröffentlicht: (2025) -
RSTGCN: Railway-centric Spatio-Temporal Graph Convolutional Network for Train Delay Prediction
von: Chowdhury, Koyena, et al.
Veröffentlicht: (2025) -
Advancing Decoding Strategies: Enhancements in Locally Typical Sampling for LLMs
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)