Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Weilin, Huang, Yuxiang, Han, Xu, Xu, Wang, Xiao, Chaojun, Zhang, Xinrong, Fang, Yewei, Zhang, Kaihuo, Liu, Zhiyuan, Sun, Maosong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling
by: Zhao, Weilin, et al.
Published: (2025)
by: Zhao, Weilin, et al.
Published: (2025)
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
by: Huang, Yuxiang, et al.
Published: (2025)
by: Huang, Yuxiang, et al.
Published: (2025)
Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding
by: Huang, Jianuo, et al.
Published: (2026)
by: Huang, Jianuo, et al.
Published: (2026)
Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models
by: Zhang, Xinrong, et al.
Published: (2024)
by: Zhang, Xinrong, et al.
Published: (2024)
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
by: Song, Chenyang, et al.
Published: (2025)
by: Song, Chenyang, et al.
Published: (2025)
Connecting the Dots: Inferring Patent Phrase Similarity with Retrieved Phrase Graphs
by: Peng, Zhuoyi, et al.
Published: (2024)
by: Peng, Zhuoyi, et al.
Published: (2024)
Cascade Speculative Drafting for Even Faster LLM Inference
by: Chen, Ziyi, et al.
Published: (2023)
by: Chen, Ziyi, et al.
Published: (2023)
Generalised Medical Phrase Grounding
by: Zhang, Wenjun, et al.
Published: (2025)
by: Zhang, Wenjun, et al.
Published: (2025)
Balancing Coverage and Draft Latency in Vocabulary Trimming for Faster Speculative Decoding
by: Shoham, Ofir Ben
Published: (2026)
by: Shoham, Ofir Ben
Published: (2026)
Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads on Consumer-Grade Devices
by: Huang, Yuxiang, et al.
Published: (2024)
by: Huang, Yuxiang, et al.
Published: (2024)
PEARL: Parallel Speculative Decoding with Adaptive Draft Length
by: Liu, Tianyu, et al.
Published: (2024)
by: Liu, Tianyu, et al.
Published: (2024)
FlexDraft: Flexible Speculative Decoding via Attention Tuning and Bonus-Guided Calibration
by: Zhang, Yaojie, et al.
Published: (2026)
by: Zhang, Yaojie, et al.
Published: (2026)
Speculative Decoding Meets Quantization: Compatibility Evaluation and Hierarchical Framework Design
by: Zhang, Yudi, et al.
Published: (2025)
by: Zhang, Yudi, et al.
Published: (2025)
Open (Clinical) LLMs are Sensitive to Instruction Phrasings
by: Arroyo, Alberto Mario Ceballos, et al.
Published: (2024)
by: Arroyo, Alberto Mario Ceballos, et al.
Published: (2024)
POSS: Position Specialist Generates Better Draft for Speculative Decoding
by: Huang, Langlin, et al.
Published: (2025)
by: Huang, Langlin, et al.
Published: (2025)
SpecBlock: Block-Iterative Speculative Decoding with Dynamic Tree Drafting
by: Shi, Weijie, et al.
Published: (2026)
by: Shi, Weijie, et al.
Published: (2026)
CA-LoRA: Adapting Existing LoRA for Compressed LLMs to Enable Efficient Multi-Tasking on Personal Devices
by: Zhao, Weilin, et al.
Published: (2023)
by: Zhao, Weilin, et al.
Published: (2023)
SJD-PV: Speculative Jacobi Decoding with Phrase Verification for Autoregressive Image Generation
by: Yu, Zhehao, et al.
Published: (2026)
by: Yu, Zhehao, et al.
Published: (2026)
Cost-Aware Diffusion Draft Trees for Speculative Decoding
by: Zhang, Shuai, et al.
Published: (2026)
by: Zhang, Shuai, et al.
Published: (2026)
Head-Driven Phrase Structure Grammar
by: Müller, Stefan, et al.
Published: (2025)
by: Müller, Stefan, et al.
Published: (2025)
Language Model as an Annotator: Unsupervised Context-aware Quality Phrase Generation
by: Zhang, Zhihao, et al.
Published: (2023)
by: Zhang, Zhihao, et al.
Published: (2023)
Learning to Draft: Adaptive Speculative Decoding with Reinforcement Learning
by: Zhang, Jiebin, et al.
Published: (2026)
by: Zhang, Jiebin, et al.
Published: (2026)
DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
by: Song, Chenyang, et al.
Published: (2026)
by: Song, Chenyang, et al.
Published: (2026)
Similar Phrases for Cause of Actions of Civil Cases
by: Huang, Ho-Chien, et al.
Published: (2024)
by: Huang, Ho-Chien, et al.
Published: (2024)
LR Parsing of Permutation Phrases
by: Kostičová, Jana
Published: (2024)
by: Kostičová, Jana
Published: (2024)
Learning High-Quality and General-Purpose Phrase Representations
by: Chen, Lihu, et al.
Published: (2024)
by: Chen, Lihu, et al.
Published: (2024)
Cross-lingual Contextualized Phrase Retrieval
by: Li, Huayang, et al.
Published: (2024)
by: Li, Huayang, et al.
Published: (2024)
Densing Law of LLMs
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
by: Zhang, Ziyin, et al.
Published: (2024)
by: Zhang, Ziyin, et al.
Published: (2024)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
Stuffed Mamba: Oversized States Lead to the Inability to Forget
by: Chen, Yingfa, et al.
Published: (2024)
by: Chen, Yingfa, et al.
Published: (2024)
SpecHub: Provable Acceleration to Multi-Draft Speculative Decoding
by: Sun, Ryan, et al.
Published: (2024)
by: Sun, Ryan, et al.
Published: (2024)
Head-Driven Phrase Structure Grammar
Published: (2022)
Published: (2022)
Phrase-Instance Alignment for Generalized Referring Segmentation
by: Nguyen, E-Ro, et al.
Published: (2024)
by: Nguyen, E-Ro, et al.
Published: (2024)
MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
by: Hu, Shengding, et al.
Published: (2024)
by: Hu, Shengding, et al.
Published: (2024)
OPT-Tree: Speculative Decoding with Adaptive Draft Tree Structure
by: Wang, Jikai, et al.
Published: (2024)
by: Wang, Jikai, et al.
Published: (2024)
General Phrase Debiaser: Debiasing Masked Language Models at a Multi-Token Level
by: Shi, Bingkang, et al.
Published: (2023)
by: Shi, Bingkang, et al.
Published: (2023)
KOALA: Enhancing Speculative Decoding for LLM via Multi-Layer Draft Heads with Adversarial Learning
by: Zhang, Kaiqi, et al.
Published: (2024)
by: Zhang, Kaiqi, et al.
Published: (2024)
$\texttt{SPECS}$: Faster Test-Time Scaling through Speculative Drafts
by: Cemri, Mert, et al.
Published: (2025)
by: Cemri, Mert, et al.
Published: (2025)
PHRASED: Phrase Dictionary Biasing for Speech Translation
by: Wang, Peidong, et al.
Published: (2025)
by: Wang, Peidong, et al.
Published: (2025)
Similar Items
-
FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling
by: Zhao, Weilin, et al.
Published: (2025) -
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
by: Huang, Yuxiang, et al.
Published: (2025) -
Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding
by: Huang, Jianuo, et al.
Published: (2026) -
Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models
by: Zhang, Xinrong, et al.
Published: (2024) -
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
by: Song, Chenyang, et al.
Published: (2025)