Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding
Fuente:
arXiv
Saved in:
| Main Authors: | Cho, Sukmin, Choi, Sangjin, Hwang, Taeho, Seo, Jeongyeon, Jeong, Soyeong, Lee, Huije, Song, Hoyun, Park, Jong C., Kwon, Youngjin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Typos that Broke the RAG's Back: Genetic Attack on RAG Pipeline by Simulating Documents in the Wild via Low-level Perturbations
by: Cho, Sukmin, et al.
Published: (2024)
by: Cho, Sukmin, et al.
Published: (2024)
Temporal Information Retrieval via Time-Specifier Model Merging
by: Han, SeungYoon, et al.
Published: (2025)
by: Han, SeungYoon, et al.
Published: (2025)
EXIT: Context-Aware Extractive Compression for Enhancing Retrieval-Augmented Generation
by: Hwang, Taeho, et al.
Published: (2024)
by: Hwang, Taeho, et al.
Published: (2024)
Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
by: Shin, Jisu, et al.
Published: (2024)
by: Shin, Jisu, et al.
Published: (2024)
DSLR: Document Refinement with Sentence-Level Re-ranking and Reconstruction to Enhance Retrieval-Augmented Generation
by: Hwang, Taeho, et al.
Published: (2024)
by: Hwang, Taeho, et al.
Published: (2024)
Does Rationale Quality Matter? Enhancing Mental Disorder Detection via Selective Reasoning Distillation
by: Song, Hoyun, et al.
Published: (2025)
by: Song, Hoyun, et al.
Published: (2025)
Towards Effective Counter-Responses: Aligning Human Preferences with Strategies to Combat Online Trolling
by: Lee, Huije, et al.
Published: (2024)
by: Lee, Huije, et al.
Published: (2024)
Self-Knowledge Distillation for Learning Ambiguity
by: Park, Hancheol, et al.
Published: (2024)
by: Park, Hancheol, et al.
Published: (2024)
A Spatio-Temporal Representation Learning as an Alternative to Traditional Glosses in Sign Language Translation and Production
by: Hwang, Eui Jun, et al.
Published: (2024)
by: Hwang, Eui Jun, et al.
Published: (2024)
Database-Augmented Query Representation for Information Retrieval
by: Jeong, Soyeong, et al.
Published: (2024)
by: Jeong, Soyeong, et al.
Published: (2024)
Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity
by: Jeong, Soyeong, et al.
Published: (2024)
by: Jeong, Soyeong, et al.
Published: (2024)
Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach
by: Ko, Changgeon, et al.
Published: (2024)
by: Ko, Changgeon, et al.
Published: (2024)
Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding
by: Zhang, Jun, et al.
Published: (2023)
by: Zhang, Jun, et al.
Published: (2023)
The RAG Paradox: A Black-Box Attack Exploiting Unintentional Vulnerabilities in Retrieval-Augmented Generation Systems
by: Choi, Chanwoo, et al.
Published: (2025)
by: Choi, Chanwoo, et al.
Published: (2025)
Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation
by: Lee, Huije, et al.
Published: (2026)
by: Lee, Huije, et al.
Published: (2026)
Social Dynamics as Critical Vulnerabilities that Undermine Objective Decision-Making in LLM Collectives
by: Ko, Changgeon, et al.
Published: (2026)
by: Ko, Changgeon, et al.
Published: (2026)
A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings
by: Gaim, Fitsum, et al.
Published: (2025)
by: Gaim, Fitsum, et al.
Published: (2025)
Upcycling Candidate Tokens of Large Language Models for Query Expansion
by: Kim, Jinseok, et al.
Published: (2025)
by: Kim, Jinseok, et al.
Published: (2025)
FastEagle: Cascaded Drafting for Accelerating Speculative Decoding
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
by: Yang, Penghui, et al.
Published: (2025)
by: Yang, Penghui, et al.
Published: (2025)
TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
by: Sun, Hanshi, et al.
Published: (2024)
by: Sun, Hanshi, et al.
Published: (2024)
Overcoming Joint Intractability with Lossless Hierarchical Speculative Decoding
by: Zhou, Yuxuan, et al.
Published: (2026)
by: Zhou, Yuxuan, et al.
Published: (2026)
Speculative Coupled Decoding for Training-Free Lossless Acceleration of Autoregressive Visual Generation
by: So, Junhyuk, et al.
Published: (2025)
by: So, Junhyuk, et al.
Published: (2025)
Accelerating Speculative Decoding with Block Diffusion Draft Trees
by: Ringel, Liran, et al.
Published: (2026)
by: Ringel, Liran, et al.
Published: (2026)
SpecHub: Provable Acceleration to Multi-Draft Speculative Decoding
by: Sun, Ryan, et al.
Published: (2024)
by: Sun, Ryan, et al.
Published: (2024)
SCANNER: Knowledge-Enhanced Approach for Robust Multi-modal Named Entity Recognition of Unseen Entities
by: Ok, Hyunjong, et al.
Published: (2024)
by: Ok, Hyunjong, et al.
Published: (2024)
PiLaMIM: Toward Richer Visual Representations by Integrating Pixel and Latent Masked Image Modeling
by: Lee, Junmyeong, et al.
Published: (2025)
by: Lee, Junmyeong, et al.
Published: (2025)
An Efficient Sign Language Translation Using Spatial Configuration and Motion Dynamics with LLMs
by: Hwang, Eui Jun, et al.
Published: (2024)
by: Hwang, Eui Jun, et al.
Published: (2024)
Autoregressive Sign Language Production: A Gloss-Free Approach with Discrete Representations
by: Hwang, Eui Jun, et al.
Published: (2023)
by: Hwang, Eui Jun, et al.
Published: (2023)
Spiffy: Multiplying Diffusion LLM Acceleration via Lossless Speculative Decoding
by: Agrawal, Sudhanshu, et al.
Published: (2025)
by: Agrawal, Sudhanshu, et al.
Published: (2025)
Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
by: Timor, Nadav, et al.
Published: (2025)
by: Timor, Nadav, et al.
Published: (2025)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
by: Wen, Zhuofan, et al.
Published: (2024)
by: Wen, Zhuofan, et al.
Published: (2024)
Speculate Deep and Accurate: Lossless and Training-Free Acceleration for Offloaded LLMs via Substitute Speculative Decoding
by: Wang, Pei-Shuo, et al.
Published: (2025)
by: Wang, Pei-Shuo, et al.
Published: (2025)
SpecTr-GBV: Multi-Draft Block Verification Accelerating Speculative Decoding
by: Lin, Yijun, et al.
Published: (2026)
by: Lin, Yijun, et al.
Published: (2026)
SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding
by: Plaksin, Anton, et al.
Published: (2026)
by: Plaksin, Anton, et al.
Published: (2026)
Learning to Draft: Adaptive Speculative Decoding with Reinforcement Learning
by: Zhang, Jiebin, et al.
Published: (2026)
by: Zhang, Jiebin, et al.
Published: (2026)
SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation
by: Kang, Jialiang, et al.
Published: (2026)
by: Kang, Jialiang, et al.
Published: (2026)
DFlare: Scaling Up Draft Capacity for Block Diffusion Speculative Decoding
by: Zhang, Jiebin, et al.
Published: (2026)
by: Zhang, Jiebin, et al.
Published: (2026)
LLM Watermark Evasion via Bias Inversion
by: Hwang, Jeongyeon, et al.
Published: (2025)
by: Hwang, Jeongyeon, et al.
Published: (2025)
CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs
by: Ning, Zhiyuan, et al.
Published: (2025)
by: Ning, Zhiyuan, et al.
Published: (2025)
Similar Items
-
Typos that Broke the RAG's Back: Genetic Attack on RAG Pipeline by Simulating Documents in the Wild via Low-level Perturbations
by: Cho, Sukmin, et al.
Published: (2024) -
Temporal Information Retrieval via Time-Specifier Model Merging
by: Han, SeungYoon, et al.
Published: (2025) -
EXIT: Context-Aware Extractive Compression for Enhancing Retrieval-Augmented Generation
by: Hwang, Taeho, et al.
Published: (2024) -
Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
by: Shin, Jisu, et al.
Published: (2024) -
DSLR: Document Refinement with Sentence-Level Re-ranking and Reconstruction to Enhance Retrieval-Augmented Generation
by: Hwang, Taeho, et al.
Published: (2024)