P-EAGLE: Parallel-Drafting EAGLE with Scalable Training
Fuente:
arXiv
Saved in:
| Main Authors: | Hui, Mude, Huang, Xin, Salas, Jaime Campos, Sun, Yue, Pemberton, Nathan, Song, Xiang, Khetan, Ashish, Karypis, George |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AdaEAGLE: Optimizing Speculative Decoding via Explicit Modeling of Adaptive Draft Structures
by: Zhang, Situo, et al.
Published: (2024)
by: Zhang, Situo, et al.
Published: (2024)
XShare: Collaborative in-Batch Expert Sharing for Faster MoE Inference
by: Vankov, Daniil, et al.
Published: (2026)
by: Vankov, Daniil, et al.
Published: (2026)
EAGLE: Egocentric AGgregated Language-video Engine
by: Bi, Jing, et al.
Published: (2024)
by: Bi, Jing, et al.
Published: (2024)
EAGLE: Contrastive Learning for Efficient Graph Anomaly Detection
by: Ren, Jing, et al.
Published: (2025)
by: Ren, Jing, et al.
Published: (2025)
When LLMs get significantly worse: A statistical approach to detect model degradations
by: Kübler, Jonas, et al.
Published: (2026)
by: Kübler, Jonas, et al.
Published: (2026)
EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models
by: Villa, Andrés, et al.
Published: (2025)
by: Villa, Andrés, et al.
Published: (2025)
EAGLE: A Domain Generalization Framework for AI-generated Text Detection
by: Bhattacharjee, Amrita, et al.
Published: (2024)
by: Bhattacharjee, Amrita, et al.
Published: (2024)
EAGLE: Efficient Alignment of Generalized Latent Embeddings for Multimodal Survival Prediction with Interpretable Attribution Analysis
by: Tripathi, Aakash, et al.
Published: (2025)
by: Tripathi, Aakash, et al.
Published: (2025)
EAGLE: Edge-Aware Graph Learning for Proactive Delivery Delay Prediction in Smart Logistics Networks
by: Xue, Zhiming, et al.
Published: (2026)
by: Xue, Zhiming, et al.
Published: (2026)
SemPool: Simple, robust, and interpretable KG pooling for enhancing language models
by: Mavromatis, Costas, et al.
Published: (2024)
by: Mavromatis, Costas, et al.
Published: (2024)
Accelerating PayPal's Commerce Agent with Speculative Decoding: An Empirical Study on EAGLE3 with Fine-Tuned Nemotron Models
by: Qin, Ally, et al.
Published: (2026)
by: Qin, Ally, et al.
Published: (2026)
PoliticsBench: Benchmarking Political Values in Large Language Models with Multi-Turn Roleplay
by: Khetan, Rohan, et al.
Published: (2026)
by: Khetan, Rohan, et al.
Published: (2026)
DEFT: Data Efficient Fine-Tuning for Pre-Trained Language Models via Unsupervised Core-Set Selection
by: Das, Devleena, et al.
Published: (2023)
by: Das, Devleena, et al.
Published: (2023)
EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees
by: Li, Yuhui, et al.
Published: (2024)
by: Li, Yuhui, et al.
Published: (2024)
GNN-RAG: Graph Neural Retrieval for Large Language Model Reasoning
by: Mavromatis, Costas, et al.
Published: (2024)
by: Mavromatis, Costas, et al.
Published: (2024)
PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models
by: Friedland, Gerald, et al.
Published: (2024)
by: Friedland, Gerald, et al.
Published: (2024)
Exploring and Improving Drafts in Blockwise Parallel Decoding
by: Kim, Taehyeon, et al.
Published: (2024)
by: Kim, Taehyeon, et al.
Published: (2024)
Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree
by: Gao, Xiangxiang, et al.
Published: (2024)
by: Gao, Xiangxiang, et al.
Published: (2024)
D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting
by: Wu, Tianyu, et al.
Published: (2026)
by: Wu, Tianyu, et al.
Published: (2026)
MixGCN: Scalable GCN Training by Mixture of Parallelism and Mixture of Accelerators
by: Wan, Cheng, et al.
Published: (2025)
by: Wan, Cheng, et al.
Published: (2025)
Scalable Prompt Routing via Fine-Grained Latent Task Discovery
by: Zhang, Yunyi, et al.
Published: (2026)
by: Zhang, Yunyi, et al.
Published: (2026)
PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding
by: An, Zihao, et al.
Published: (2026)
by: An, Zihao, et al.
Published: (2026)
Free Draft-and-Verification: Toward Lossless Parallel Decoding for Diffusion Large Language Models
by: Wu, Shutong, et al.
Published: (2025)
by: Wu, Shutong, et al.
Published: (2025)
DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verification, and Fully Parallel Execution
by: Hu, Yunhai, et al.
Published: (2026)
by: Hu, Yunhai, et al.
Published: (2026)
Ultra-diffuse galaxies in the EAGLE simulation
by: Zheng, Haonan, et al.
Published: (2025)
by: Zheng, Haonan, et al.
Published: (2025)
TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training
by: Liu, Man, et al.
Published: (2026)
by: Liu, Man, et al.
Published: (2026)
Communication-free Sampling and 4D Hybrid Parallelism for Scalable Mini-batch GNN Training
by: Wei, Cunyang, et al.
Published: (2026)
by: Wei, Cunyang, et al.
Published: (2026)
Scalable and Adaptive Parallel Training of Graph Transformer on Large Graphs
by: Lin, Jun-Liang, et al.
Published: (2026)
by: Lin, Jun-Liang, et al.
Published: (2026)
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
by: Shen, Yuhao, et al.
Published: (2025)
by: Shen, Yuhao, et al.
Published: (2025)
MineDraft: A Framework for Batch Parallel Speculative Decoding
by: Tang, Zhenwei, et al.
Published: (2026)
by: Tang, Zhenwei, et al.
Published: (2026)
Draft-and-Prune: Improving the Reliability of Auto-formalization for Logical Reasoning
by: Ni, Zhiyu, et al.
Published: (2026)
by: Ni, Zhiyu, et al.
Published: (2026)
SpecHub: Provable Acceleration to Multi-Draft Speculative Decoding
by: Sun, Ryan, et al.
Published: (2024)
by: Sun, Ryan, et al.
Published: (2024)
Hierarchical Compression of Text-Rich Graphs via Large Language Models
by: Zhang, Shichang, et al.
Published: (2024)
by: Zhang, Shichang, et al.
Published: (2024)
Legal Documents Drafting with Fine-Tuned Pre-Trained Large Language Model
by: Lin, Chun-Hsien, et al.
Published: (2024)
by: Lin, Chun-Hsien, et al.
Published: (2024)
AEC — Axiomatic Engine Cycle v0.1 — Initial Specification Draft
by: Brown, Cameron
Published: (2026)
by: Brown, Cameron
Published: (2026)
TAPS: Target-Aware Prefix Tree Selection for Diffusion-Drafted Speculative Decoding
by: Wang, Zhuoyu, et al.
Published: (2026)
by: Wang, Zhuoyu, et al.
Published: (2026)
From Demonstrations to Rewards: Alignment Without Explicit Human Preferences
by: Zeng, Siliang, et al.
Published: (2025)
by: Zeng, Siliang, et al.
Published: (2025)
InfoSeeker: A Scalable Hierarchical Parallel Agent Framework for Web Information Seeking
by: Lee, Ka Yiu, et al.
Published: (2026)
by: Lee, Ka Yiu, et al.
Published: (2026)
ODRL Policy Comparison Through Normalisation
by: Salas, Jaime Osvaldo, et al.
Published: (2026)
by: Salas, Jaime Osvaldo, et al.
Published: (2026)
Hybrid Deep Searcher: Scalable Parallel and Sequential Search Reasoning
by: Ko, Dayoon, et al.
Published: (2025)
by: Ko, Dayoon, et al.
Published: (2025)
Similar Items
-
AdaEAGLE: Optimizing Speculative Decoding via Explicit Modeling of Adaptive Draft Structures
by: Zhang, Situo, et al.
Published: (2024) -
XShare: Collaborative in-Batch Expert Sharing for Faster MoE Inference
by: Vankov, Daniil, et al.
Published: (2026) -
EAGLE: Egocentric AGgregated Language-video Engine
by: Bi, Jing, et al.
Published: (2024) -
EAGLE: Contrastive Learning for Efficient Graph Anomaly Detection
by: Ren, Jing, et al.
Published: (2025) -
When LLMs get significantly worse: A statistical approach to detect model degradations
by: Kübler, Jonas, et al.
Published: (2026)