S2D: Sorted Speculative Decoding For More Efficient Deployment of Nested Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kavehzadeh, Parsa, Pourreza, Mohammadreza, Valipour, Mojtaba, Zhu, Tinashu, Bai, Haoli, Ghodsi, Ali, Chen, Boxing, Rezagholizadeh, Mehdi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sorted LLaMA: Unlocking the Potential of Intermediate Layers of Large Language Models for Dynamic Inference
von: Kavehzadeh, Parsa, et al.
Veröffentlicht: (2023)
von: Kavehzadeh, Parsa, et al.
Veröffentlicht: (2023)
SortedNet: A Scalable and Generalized Framework for Training Modular Deep Neural Networks
von: Valipour, Mojtaba, et al.
Veröffentlicht: (2023)
von: Valipour, Mojtaba, et al.
Veröffentlicht: (2023)
QDyLoRA: Quantized Dynamic Low-Rank Adaptation for Efficient Large Language Model Tuning
von: Rajabzadeh, Hossein, et al.
Veröffentlicht: (2024)
von: Rajabzadeh, Hossein, et al.
Veröffentlicht: (2024)
Balcony: A Lightweight Approach to Dynamic Inference of Generative Language Models
von: Jamialahmadi, Benyamin, et al.
Veröffentlicht: (2025)
von: Jamialahmadi, Benyamin, et al.
Veröffentlicht: (2025)
EchoAtt: Attend, Copy, then Adjust for More Efficient Large Language Models
von: Rajabzadeh, Hossein, et al.
Veröffentlicht: (2024)
von: Rajabzadeh, Hossein, et al.
Veröffentlicht: (2024)
Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity
von: Metel, Michael R., et al.
Veröffentlicht: (2024)
von: Metel, Michael R., et al.
Veröffentlicht: (2024)
On the importance of Data Scale in Pretraining Arabic Language Models
von: Ghaddar, Abbas, et al.
Veröffentlicht: (2024)
von: Ghaddar, Abbas, et al.
Veröffentlicht: (2024)
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression
von: Metel, Michael R., et al.
Veröffentlicht: (2024)
von: Metel, Michael R., et al.
Veröffentlicht: (2024)
DTS-SQL: Decomposed Text-to-SQL with Small Large Language Models
von: Pourreza, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Pourreza, Mohammadreza, et al.
Veröffentlicht: (2024)
Confidence Estimation for Text-to-SQL in Large Language Models
von: Maleki, Sepideh Entezari, et al.
Veröffentlicht: (2025)
von: Maleki, Sepideh Entezari, et al.
Veröffentlicht: (2025)
Beyond the Limits: A Survey of Techniques to Extend the Context Length in Large Language Models
von: Wang, Xindi, et al.
Veröffentlicht: (2024)
von: Wang, Xindi, et al.
Veröffentlicht: (2024)
ReGLA: Refining Gated Linear Attention
von: Lu, Peng, et al.
Veröffentlicht: (2025)
von: Lu, Peng, et al.
Veröffentlicht: (2025)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
Overcoming stretching and shortening assumptions in Euler-Bernoulli theory using nonlinear Hencky beam models: applicable to partly-shortened and partly-stretched beams
von: Rezaei, Mohammad Parsa, et al.
Veröffentlicht: (2024)
von: Rezaei, Mohammad Parsa, et al.
Veröffentlicht: (2024)
DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers
von: Sharma, Aman, et al.
Veröffentlicht: (2025)
von: Sharma, Aman, et al.
Veröffentlicht: (2025)
Efficient Adaptive Rejection Sampling for Accelerating Speculative Decoding in Large Language Models
von: Sun, Chendong, et al.
Veröffentlicht: (2025)
von: Sun, Chendong, et al.
Veröffentlicht: (2025)
An Efficient End-to-End Approach to Noise Invariant Speech Features via Multi-Task Learning
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2024)
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2024)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
von: Heuillet, Maxime, et al.
Veröffentlicht: (2025)
von: Heuillet, Maxime, et al.
Veröffentlicht: (2025)
OTTAWA: Optimal TransporT Adaptive Word Aligner for Hallucination and Omission Translation Errors Detection
von: Huang, Chenyang, et al.
Veröffentlicht: (2024)
von: Huang, Chenyang, et al.
Veröffentlicht: (2024)
CHARP: Conversation History AwaReness Probing for Knowledge-grounded Dialogue Systems
von: Ghaddar, Abbas, et al.
Veröffentlicht: (2024)
von: Ghaddar, Abbas, et al.
Veröffentlicht: (2024)
Sentinel-2 for Crop Yield Estimation: A Systematic Review
von: Narimani, Mohammadreza, et al.
Veröffentlicht: (2026)
von: Narimani, Mohammadreza, et al.
Veröffentlicht: (2026)
Symbolic-Diffusion: Deep Learning Based Symbolic Regression with D3PM Discrete Token Diffusion
von: Tymkow, Ryan T., et al.
Veröffentlicht: (2025)
von: Tymkow, Ryan T., et al.
Veröffentlicht: (2025)
Mapping Tomato Cropping Systems in California Using AlphaEarth Geospatial Embeddings and Deep Learning Analysis
von: Narimani, Mohammadreza, et al.
Veröffentlicht: (2026)
von: Narimani, Mohammadreza, et al.
Veröffentlicht: (2026)
HADES: Hardware Accelerated Decoding for Efficient Speculation in Large Language Models
von: Yang, Ze, et al.
Veröffentlicht: (2024)
von: Yang, Ze, et al.
Veröffentlicht: (2024)
Dynamic Speculation Lookahead Accelerates Speculative Decoding of Large Language Models
von: Mamou, Jonathan, et al.
Veröffentlicht: (2024)
von: Mamou, Jonathan, et al.
Veröffentlicht: (2024)
Resonance RoPE: Improving Context Length Generalization of Large Language Models
von: Wang, Suyuchen, et al.
Veröffentlicht: (2024)
von: Wang, Suyuchen, et al.
Veröffentlicht: (2024)
On Speculative Decoding for Multimodal Large Language Models
von: Gagrani, Mukul, et al.
Veröffentlicht: (2024)
von: Gagrani, Mukul, et al.
Veröffentlicht: (2024)
CHIQ: Contextual History Enhancement for Improving Query Rewriting in Conversational Search
von: Mo, Fengran, et al.
Veröffentlicht: (2024)
von: Mo, Fengran, et al.
Veröffentlicht: (2024)
How Many Heads Make an SSM? A Unified Framework for Attention and State Space Models
von: Ghodsi, Ali
Veröffentlicht: (2025)
von: Ghodsi, Ali
Veröffentlicht: (2025)
NeRCC: Nested-Regression Coded Computing for Resilient Distributed Prediction Serving Systems
von: Moradi, Parsa, et al.
Veröffentlicht: (2024)
von: Moradi, Parsa, et al.
Veröffentlicht: (2024)
DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking
von: Haridas, Akash, et al.
Veröffentlicht: (2026)
von: Haridas, Akash, et al.
Veröffentlicht: (2026)
CHESS: Contextual Harnessing for Efficient SQL Synthesis
von: Talaei, Shayan, et al.
Veröffentlicht: (2024)
von: Talaei, Shayan, et al.
Veröffentlicht: (2024)
Self Speculative Decoding for Diffusion Large Language Models
von: Gao, Yifeng, et al.
Veröffentlicht: (2025)
von: Gao, Yifeng, et al.
Veröffentlicht: (2025)
Confidence-Modulated Speculative Decoding for Large Language Models
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
Speculative Decoding Reimagined for Multimodal Large Language Models
von: Lin, Luxi, et al.
Veröffentlicht: (2025)
von: Lin, Luxi, et al.
Veröffentlicht: (2025)
EMS-SD: Efficient Multi-sample Speculative Decoding for Accelerating Large Language Models
von: Ni, Yunsheng, et al.
Veröffentlicht: (2024)
von: Ni, Yunsheng, et al.
Veröffentlicht: (2024)
DSSD: Efficient Edge-Device LLM Deployment and Collaborative Inference via Distributed Split Speculative Decoding
von: Ning, Jiahong, et al.
Veröffentlicht: (2025)
von: Ning, Jiahong, et al.
Veröffentlicht: (2025)
Speculative Decoding Across Languages
von: Paudel, Nirajan, et al.
Veröffentlicht: (2026)
von: Paudel, Nirajan, et al.
Veröffentlicht: (2026)
Speculative Speculative Decoding
von: Kumar, Tanishq, et al.
Veröffentlicht: (2026)
von: Kumar, Tanishq, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Sorted LLaMA: Unlocking the Potential of Intermediate Layers of Large Language Models for Dynamic Inference
von: Kavehzadeh, Parsa, et al.
Veröffentlicht: (2023) -
SortedNet: A Scalable and Generalized Framework for Training Modular Deep Neural Networks
von: Valipour, Mojtaba, et al.
Veröffentlicht: (2023) -
QDyLoRA: Quantized Dynamic Low-Rank Adaptation for Efficient Large Language Model Tuning
von: Rajabzadeh, Hossein, et al.
Veröffentlicht: (2024) -
Balcony: A Lightweight Approach to Dynamic Inference of Generative Language Models
von: Jamialahmadi, Benyamin, et al.
Veröffentlicht: (2025) -
EchoAtt: Attend, Copy, then Adjust for More Efficient Large Language Models
von: Rajabzadeh, Hossein, et al.
Veröffentlicht: (2024)