SpecExtend: A Drop-in Enhancement for Speculative Decoding of Long Sequences
Fuente:
arXiv
Salvato in:
| Autori principali: | Cha, Jungyoub, Kim, Hyunjong, Cho, Sungzoon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CopySpec: Accelerating LLMs with Speculative Copy-and-Paste Without Compromising Quality
di: Dumitru, Razvan-Gabriel, et al.
Pubblicazione: (2025)
di: Dumitru, Razvan-Gabriel, et al.
Pubblicazione: (2025)
SSSD: Simply-Scalable Speculative Decoding
di: Marzollo, Michele, et al.
Pubblicazione: (2024)
di: Marzollo, Michele, et al.
Pubblicazione: (2024)
SAM Decoding: Speculative Decoding via Suffix Automaton
di: Hu, Yuxuan, et al.
Pubblicazione: (2024)
di: Hu, Yuxuan, et al.
Pubblicazione: (2024)
Component-Aware Self-Speculative Decoding in Hybrid Language Models
di: Borobia, Hector, et al.
Pubblicazione: (2026)
di: Borobia, Hector, et al.
Pubblicazione: (2026)
Attention Drift: What Autoregressive Speculative Decoding Models Learn
di: Eldenk, Doğaç, et al.
Pubblicazione: (2026)
di: Eldenk, Doğaç, et al.
Pubblicazione: (2026)
DIVERSED: Relaxed Speculative Decoding via Dynamic Ensemble Verification
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
FED-FSTQ: Fisher-Guided Token Quantization for Communication-Efficient Federated Fine-Tuning of LLMs on Edge Devices
di: Li, Changyu, et al.
Pubblicazione: (2026)
di: Li, Changyu, et al.
Pubblicazione: (2026)
Robustness of Large Language Models to Perturbations in Text
di: Singh, Ayush, et al.
Pubblicazione: (2024)
di: Singh, Ayush, et al.
Pubblicazione: (2024)
Sycophancy as compositions of Atomic Psychometric Traits
di: Jain, Shreyans, et al.
Pubblicazione: (2025)
di: Jain, Shreyans, et al.
Pubblicazione: (2025)
DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving
di: Yang, Mingyu, et al.
Pubblicazione: (2025)
di: Yang, Mingyu, et al.
Pubblicazione: (2025)
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
di: Li, Xiangchen, et al.
Pubblicazione: (2026)
di: Li, Xiangchen, et al.
Pubblicazione: (2026)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
Speculative Automated Refactoring of Imperative Deep Learning Programs to Graph Execution
di: Khatchadourian, Raffi, et al.
Pubblicazione: (2025)
di: Khatchadourian, Raffi, et al.
Pubblicazione: (2025)
ADALog: Adaptive Unsupervised Anomaly detection in Logs with Self-attention Masked Language Model
di: Pospieszny, Przemek, et al.
Pubblicazione: (2025)
di: Pospieszny, Przemek, et al.
Pubblicazione: (2025)
BERTologyNavigator: Advanced Question Answering with BERT-based Semantics
di: Rajpal, Shreya, et al.
Pubblicazione: (2024)
di: Rajpal, Shreya, et al.
Pubblicazione: (2024)
Evaluating Class Membership Relations in Knowledge Graphs using Large Language Models
di: Allen, Bradley P., et al.
Pubblicazione: (2024)
di: Allen, Bradley P., et al.
Pubblicazione: (2024)
Large Language Models Can Better Understand Knowledge Graphs Than We Thought
di: Dai, Xinbang, et al.
Pubblicazione: (2024)
di: Dai, Xinbang, et al.
Pubblicazione: (2024)
Pre-trained Language Model with Prompts for Temporal Knowledge Graph Completion
di: Xu, Wenjie, et al.
Pubblicazione: (2023)
di: Xu, Wenjie, et al.
Pubblicazione: (2023)
Question Answering Over Spatio-Temporal Knowledge Graph
di: Dai, Xinbang, et al.
Pubblicazione: (2024)
di: Dai, Xinbang, et al.
Pubblicazione: (2024)
Breaking Free Transformer Models: Task-specific Context Attribution Promises Improved Generalizability Without Fine-tuning Pre-trained LLMs
di: Tytarenko, Stepan, et al.
Pubblicazione: (2024)
di: Tytarenko, Stepan, et al.
Pubblicazione: (2024)
ROZA Graphs: Self-Improving Near-Deterministic RAG through Evidence-Centric Feedback
di: Penaroza, Matthew
Pubblicazione: (2026)
di: Penaroza, Matthew
Pubblicazione: (2026)
KNOW: A Real-World Ontology for Knowledge Capture with Large Language Models
di: Bendiken, Arto
Pubblicazione: (2024)
di: Bendiken, Arto
Pubblicazione: (2024)
FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft Learning
di: Zhang, Yizhou, et al.
Pubblicazione: (2025)
di: Zhang, Yizhou, et al.
Pubblicazione: (2025)
Unveiling Transformer Perception by Exploring Input Manifolds
di: Benfenati, Alessandro, et al.
Pubblicazione: (2024)
di: Benfenati, Alessandro, et al.
Pubblicazione: (2024)
Generative AI for Strategic Plan Development
di: Ponnock, Jesse
Pubblicazione: (2025)
di: Ponnock, Jesse
Pubblicazione: (2025)
GPT-4 Generated Narratives of Life Events using a Structured Narrative Prompt: A Validation Study
di: Lynch, Christopher J., et al.
Pubblicazione: (2024)
di: Lynch, Christopher J., et al.
Pubblicazione: (2024)
Encoder vs Decoder: Comparative Analysis of Encoder and Decoder Language Models on Multilingual NLU Tasks
di: Nielsen, Dan Saattrup, et al.
Pubblicazione: (2024)
di: Nielsen, Dan Saattrup, et al.
Pubblicazione: (2024)
Efficient Adaptive Rejection Sampling for Accelerating Speculative Decoding in Large Language Models
di: Sun, Chendong, et al.
Pubblicazione: (2025)
di: Sun, Chendong, et al.
Pubblicazione: (2025)
Generalizing Numerical Reasoning in Table Data through Operation Sketches and Self-Supervised Learning
di: Cho, Hanjun, et al.
Pubblicazione: (2026)
di: Cho, Hanjun, et al.
Pubblicazione: (2026)
A Pluggable Common Sense-Enhanced Framework for Knowledge Graph Completion
di: Niu, Guanglin, et al.
Pubblicazione: (2024)
di: Niu, Guanglin, et al.
Pubblicazione: (2024)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
di: Saji, Alan, et al.
Pubblicazione: (2025)
di: Saji, Alan, et al.
Pubblicazione: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
di: Peters, Sydney, et al.
Pubblicazione: (2025)
di: Peters, Sydney, et al.
Pubblicazione: (2025)
Graph Language Models
di: Plenz, Moritz, et al.
Pubblicazione: (2024)
di: Plenz, Moritz, et al.
Pubblicazione: (2024)
RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale
di: Goldstein, Daniel, et al.
Pubblicazione: (2025)
di: Goldstein, Daniel, et al.
Pubblicazione: (2025)
Dodo: Dynamic Contextual Compression for Decoder-only LMs
di: Qin, Guanghui, et al.
Pubblicazione: (2023)
di: Qin, Guanghui, et al.
Pubblicazione: (2023)
LongDocFACTScore: Evaluating the Factuality of Long Document Abstractive Summarisation
di: Bishop, Jennifer A, et al.
Pubblicazione: (2023)
di: Bishop, Jennifer A, et al.
Pubblicazione: (2023)
The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary
di: Guo, Dongxin, et al.
Pubblicazione: (2026)
di: Guo, Dongxin, et al.
Pubblicazione: (2026)
Fair-PP: A Synthetic Dataset for Aligning LLM with Personalized Preferences of Social Equity
di: Zhou, Qi, et al.
Pubblicazione: (2025)
di: Zhou, Qi, et al.
Pubblicazione: (2025)
Large Language Model (LLM) Bias Index -- LLMBI
di: Oketunji, Abiodun Finbarrs, et al.
Pubblicazione: (2023)
di: Oketunji, Abiodun Finbarrs, et al.
Pubblicazione: (2023)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
di: Cho, Seonglae, et al.
Pubblicazione: (2026)
di: Cho, Seonglae, et al.
Pubblicazione: (2026)
Documenti analoghi
-
CopySpec: Accelerating LLMs with Speculative Copy-and-Paste Without Compromising Quality
di: Dumitru, Razvan-Gabriel, et al.
Pubblicazione: (2025) -
SSSD: Simply-Scalable Speculative Decoding
di: Marzollo, Michele, et al.
Pubblicazione: (2024) -
SAM Decoding: Speculative Decoding via Suffix Automaton
di: Hu, Yuxuan, et al.
Pubblicazione: (2024) -
Component-Aware Self-Speculative Decoding in Hybrid Language Models
di: Borobia, Hector, et al.
Pubblicazione: (2026) -
Attention Drift: What Autoregressive Speculative Decoding Models Learn
di: Eldenk, Doğaç, et al.
Pubblicazione: (2026)