Salvato in:
| Autori principali: | Hoang, Duc, Jaiswal, Ajay, Samragh, Mohammad, Cho, Minsik |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2602.03921 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
di: Kim, Han-Byul, et al.
Pubblicazione: (2025)
di: Kim, Han-Byul, et al.
Pubblicazione: (2025)
TIDE: Every Layer Knows the Token Beneath the Context
di: Jaiswal, Ajay, et al.
Pubblicazione: (2026)
di: Jaiswal, Ajay, et al.
Pubblicazione: (2026)
MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
di: Zibakhsh, Soheil, et al.
Pubblicazione: (2025)
di: Zibakhsh, Soheil, et al.
Pubblicazione: (2025)
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
di: Armandpour, Mohammadreza, et al.
Pubblicazione: (2026)
di: Armandpour, Mohammadreza, et al.
Pubblicazione: (2026)
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
di: Hannah, Lauren. A, et al.
Pubblicazione: (2025)
di: Hannah, Lauren. A, et al.
Pubblicazione: (2025)
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding
di: Bang, Jehyeon, et al.
Pubblicazione: (2026)
di: Bang, Jehyeon, et al.
Pubblicazione: (2026)
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
di: Samragh, Mohammad, et al.
Pubblicazione: (2024)
di: Samragh, Mohammad, et al.
Pubblicazione: (2024)
Finding Fantastic Experts in MoEs: A Unified Study for Expert Dropping Strategies and Observations
di: Jaiswal, Ajay, et al.
Pubblicazione: (2025)
di: Jaiswal, Ajay, et al.
Pubblicazione: (2025)
SpecExtend: A Drop-in Enhancement for Speculative Decoding of Long Sequences
di: Cha, Jungyoub, et al.
Pubblicazione: (2025)
di: Cha, Jungyoub, et al.
Pubblicazione: (2025)
HiSpec: Hierarchical Speculative Decoding for LLMs
di: Kumar, Avinash, et al.
Pubblicazione: (2025)
di: Kumar, Avinash, et al.
Pubblicazione: (2025)
BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
di: Hou, Yunlong, et al.
Pubblicazione: (2025)
di: Hou, Yunlong, et al.
Pubblicazione: (2025)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
di: Tiwari, Rishabh, et al.
Pubblicazione: (2025)
di: Tiwari, Rishabh, et al.
Pubblicazione: (2025)
MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for Transformers
di: Jaiswal, Ajay, et al.
Pubblicazione: (2026)
di: Jaiswal, Ajay, et al.
Pubblicazione: (2026)
SpecMemo: Speculative Decoding is in Your Pocket
di: Yildirim, Selin, et al.
Pubblicazione: (2025)
di: Yildirim, Selin, et al.
Pubblicazione: (2025)
Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance
di: Wang, Songsheng, et al.
Pubblicazione: (2025)
di: Wang, Songsheng, et al.
Pubblicazione: (2025)
SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
di: Pan, Rui, et al.
Pubblicazione: (2025)
di: Pan, Rui, et al.
Pubblicazione: (2025)
Speculating Experts Accelerates Inference for Mixture-of-Experts
di: Madan, Vivan, et al.
Pubblicazione: (2026)
di: Madan, Vivan, et al.
Pubblicazione: (2026)
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
di: Zhou, Yongchao, et al.
Pubblicazione: (2023)
di: Zhou, Yongchao, et al.
Pubblicazione: (2023)
Towards Low-bit Communication for Tensor Parallel LLM Inference
di: Dong, Harry, et al.
Pubblicazione: (2024)
di: Dong, Harry, et al.
Pubblicazione: (2024)
ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts
di: Georganas, Evangelos, et al.
Pubblicazione: (2025)
di: Georganas, Evangelos, et al.
Pubblicazione: (2025)
SpecExit: Accelerating Large Reasoning Model via Speculative Exit
di: Yang, Rubing, et al.
Pubblicazione: (2025)
di: Yang, Rubing, et al.
Pubblicazione: (2025)
SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths
di: Huang, Kaixuan, et al.
Pubblicazione: (2024)
di: Huang, Kaixuan, et al.
Pubblicazione: (2024)
BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning
di: Xu, Yuhang, et al.
Pubblicazione: (2026)
di: Xu, Yuhang, et al.
Pubblicazione: (2026)
KnapSpec: Self-Speculative Decoding via Adaptive Layer Selection as a Knapsack Problem
di: Cha, Seongjin, et al.
Pubblicazione: (2026)
di: Cha, Seongjin, et al.
Pubblicazione: (2026)
CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs
di: Ning, Zhiyuan, et al.
Pubblicazione: (2025)
di: Ning, Zhiyuan, et al.
Pubblicazione: (2025)
SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative Decoding
di: Li, Shenggui, et al.
Pubblicazione: (2026)
di: Li, Shenggui, et al.
Pubblicazione: (2026)
SpecBound: Adaptive Bounded Self-Speculation with Layer-wise Confidence Calibration
di: Wen, Zhuofan, et al.
Pubblicazione: (2026)
di: Wen, Zhuofan, et al.
Pubblicazione: (2026)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
di: Yang, Penghui, et al.
Pubblicazione: (2025)
di: Yang, Penghui, et al.
Pubblicazione: (2025)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
di: Fu, Qichen, et al.
Pubblicazione: (2024)
di: Fu, Qichen, et al.
Pubblicazione: (2024)
Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context
di: Alizadeh, Keivan, et al.
Pubblicazione: (2026)
di: Alizadeh, Keivan, et al.
Pubblicazione: (2026)
DynaSpec: Context-aware Dynamic Speculative Sampling for Large-Vocabulary Language Models
di: Zhang, Jinbin, et al.
Pubblicazione: (2025)
di: Zhang, Jinbin, et al.
Pubblicazione: (2025)
Do Compressed LLMs Forget Knowledge? An Experimental Study with Practical Implications
di: Hoang, Duc N. M, et al.
Pubblicazione: (2023)
di: Hoang, Duc N. M, et al.
Pubblicazione: (2023)
FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling
di: Zhao, Weilin, et al.
Pubblicazione: (2025)
di: Zhao, Weilin, et al.
Pubblicazione: (2025)
Barriers for Learning in an Evolving World: Mathematical Understanding of Loss of Plasticity
di: Joudaki, Amir, et al.
Pubblicazione: (2025)
di: Joudaki, Amir, et al.
Pubblicazione: (2025)
Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
di: Yin, Lu, et al.
Pubblicazione: (2023)
di: Yin, Lu, et al.
Pubblicazione: (2023)
LLaGA: Large Language and Graph Assistant
di: Chen, Runjin, et al.
Pubblicazione: (2024)
di: Chen, Runjin, et al.
Pubblicazione: (2024)
Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential
di: Samragh, Mohammad, et al.
Pubblicazione: (2025)
di: Samragh, Mohammad, et al.
Pubblicazione: (2025)
Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching
di: Dong, Yanhao, et al.
Pubblicazione: (2025)
di: Dong, Yanhao, et al.
Pubblicazione: (2025)
Explainable AI in Time-Sensitive Scenarios: Prefetched Offline Explanation Model
di: Russo, Fabio Michele, et al.
Pubblicazione: (2025)
di: Russo, Fabio Michele, et al.
Pubblicazione: (2025)
SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection
di: Shukla, Shikhar
Pubblicazione: (2026)
di: Shukla, Shikhar
Pubblicazione: (2026)
Documenti analoghi
-
SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
di: Kim, Han-Byul, et al.
Pubblicazione: (2025) -
TIDE: Every Layer Knows the Token Beneath the Context
di: Jaiswal, Ajay, et al.
Pubblicazione: (2026) -
MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
di: Zibakhsh, Soheil, et al.
Pubblicazione: (2025) -
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
di: Armandpour, Mohammadreza, et al.
Pubblicazione: (2026) -
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
di: Hannah, Lauren. A, et al.
Pubblicazione: (2025)