P-EAGLE: Parallel-Drafting EAGLE with Scalable Training
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866918318583054336 |
|---|---|
| author | Hui, Mude Huang, Xin Salas, Jaime Campos Sun, Yue Pemberton, Nathan Song, Xiang Khetan, Ashish Karypis, George |
| author_facet | Hui, Mude Huang, Xin Salas, Jaime Campos Sun, Yue Pemberton, Nathan Song, Xiang Khetan, Ashish Karypis, George |
| contents | Reasoning LLMs produce longer outputs, requiring speculative decoding drafters trained on extended sequences. Parallel drafting - predicting multiple tokens per forward pass - offers latency benefits over sequential generation, but training complexity scales quadratically with the product of sequence length and parallel positions, rendering long-context training impractical. We present P(arallel)-EAGLE, which transforms EAGLE from autoregressive to parallel multi-token prediction via a learnable shared hidden state. To scale training to long contexts, we develop a framework featuring attention mask pre-computation and sequence partitioning techniques, enabling gradient accumulation within individual sequences for parallel-prediction training. We implement P-EAGLE in vLLM and demonstrate speedups of 1.10-1.36x over autoregressive EAGLE-3 across GPT-OSS 120B, 20B, and Qwen3-Coder 30B. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_01469 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | P-EAGLE: Parallel-Drafting EAGLE with Scalable Training Hui, Mude Huang, Xin Salas, Jaime Campos Sun, Yue Pemberton, Nathan Song, Xiang Khetan, Ashish Karypis, George Machine Learning Artificial Intelligence Reasoning LLMs produce longer outputs, requiring speculative decoding drafters trained on extended sequences. Parallel drafting - predicting multiple tokens per forward pass - offers latency benefits over sequential generation, but training complexity scales quadratically with the product of sequence length and parallel positions, rendering long-context training impractical. We present P(arallel)-EAGLE, which transforms EAGLE from autoregressive to parallel multi-token prediction via a learnable shared hidden state. To scale training to long contexts, we develop a framework featuring attention mask pre-computation and sequence partitioning techniques, enabling gradient accumulation within individual sequences for parallel-prediction training. We implement P-EAGLE in vLLM and demonstrate speedups of 1.10-1.36x over autoregressive EAGLE-3 across GPT-OSS 120B, 20B, and Qwen3-Coder 30B. |
| title | P-EAGLE: Parallel-Drafting EAGLE with Scalable Training |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2602.01469 |