P-EAGLE: Parallel-Drafting EAGLE with Scalable Training

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hui, Mude, Huang, Xin, Salas, Jaime Campos, Sun, Yue, Pemberton, Nathan, Song, Xiang, Khetan, Ashish, Karypis, George
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918318583054336
author Hui, Mude
Huang, Xin
Salas, Jaime Campos
Sun, Yue
Pemberton, Nathan
Song, Xiang
Khetan, Ashish
Karypis, George
author_facet Hui, Mude
Huang, Xin
Salas, Jaime Campos
Sun, Yue
Pemberton, Nathan
Song, Xiang
Khetan, Ashish
Karypis, George
contents Reasoning LLMs produce longer outputs, requiring speculative decoding drafters trained on extended sequences. Parallel drafting - predicting multiple tokens per forward pass - offers latency benefits over sequential generation, but training complexity scales quadratically with the product of sequence length and parallel positions, rendering long-context training impractical. We present P(arallel)-EAGLE, which transforms EAGLE from autoregressive to parallel multi-token prediction via a learnable shared hidden state. To scale training to long contexts, we develop a framework featuring attention mask pre-computation and sequence partitioning techniques, enabling gradient accumulation within individual sequences for parallel-prediction training. We implement P-EAGLE in vLLM and demonstrate speedups of 1.10-1.36x over autoregressive EAGLE-3 across GPT-OSS 120B, 20B, and Qwen3-Coder 30B.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01469
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle P-EAGLE: Parallel-Drafting EAGLE with Scalable Training
Hui, Mude
Huang, Xin
Salas, Jaime Campos
Sun, Yue
Pemberton, Nathan
Song, Xiang
Khetan, Ashish
Karypis, George
Machine Learning
Artificial Intelligence
Reasoning LLMs produce longer outputs, requiring speculative decoding drafters trained on extended sequences. Parallel drafting - predicting multiple tokens per forward pass - offers latency benefits over sequential generation, but training complexity scales quadratically with the product of sequence length and parallel positions, rendering long-context training impractical. We present P(arallel)-EAGLE, which transforms EAGLE from autoregressive to parallel multi-token prediction via a learnable shared hidden state. To scale training to long contexts, we develop a framework featuring attention mask pre-computation and sequence partitioning techniques, enabling gradient accumulation within individual sequences for parallel-prediction training. We implement P-EAGLE in vLLM and demonstrate speedups of 1.10-1.36x over autoregressive EAGLE-3 across GPT-OSS 120B, 20B, and Qwen3-Coder 30B.
title P-EAGLE: Parallel-Drafting EAGLE with Scalable Training
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.01469