Gespeichert in:
| Hauptverfasser: | Wang, Tuowei, Li, Fengzu, Sun, Yanfan, Gao, Wei, Ren, Ju |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2605.16786 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DynaKV: Enabling Accurate and Efficient Long-Sequence LLM Decoding on Smartphones
von: Wang, Tuowei, et al.
Veröffentlicht: (2025)
von: Wang, Tuowei, et al.
Veröffentlicht: (2025)
Neuralink: Fast LLM Inference on Smartphones with Neuron Co-Activation Linking
von: Wang, Tuowei, et al.
Veröffentlicht: (2024)
von: Wang, Tuowei, et al.
Veröffentlicht: (2024)
Collaborative Speculative Inference for Efficient LLM Inference Serving
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
Cascade Speculative Drafting for Even Faster LLM Inference
von: Chen, Ziyi, et al.
Veröffentlicht: (2023)
von: Chen, Ziyi, et al.
Veröffentlicht: (2023)
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2025)
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2025)
Lever: Inference-Time Policy Reuse under Support Constraints
von: Vitenko, Ihor, et al.
Veröffentlicht: (2026)
von: Vitenko, Ihor, et al.
Veröffentlicht: (2026)
SPIRe: Boosting LLM Inference Throughput with Speculative Decoding
von: Neelam, Sanjit, et al.
Veröffentlicht: (2025)
von: Neelam, Sanjit, et al.
Veröffentlicht: (2025)
SmartBench: Is Your LLM Truly a Good Chinese Smartphone Assistant?
von: Lu, Xudong, et al.
Veröffentlicht: (2025)
von: Lu, Xudong, et al.
Veröffentlicht: (2025)
Long Exposure: Accelerating Parameter-Efficient Fine-Tuning for LLMs under Shadowy Sparsity
von: Wang, Tuowei, et al.
Veröffentlicht: (2025)
von: Wang, Tuowei, et al.
Veröffentlicht: (2025)
CATS: Cascaded Adaptive Tree Speculation for Memory-Limited LLM Inference Acceleration
von: Han, Yuning, et al.
Veröffentlicht: (2026)
von: Han, Yuning, et al.
Veröffentlicht: (2026)
DataSciBench: An LLM Agent Benchmark for Data Science
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
von: Hao, Zixu, et al.
Veröffentlicht: (2025)
von: Hao, Zixu, et al.
Veröffentlicht: (2025)
SpecPipe: Accelerating Pipeline Parallelism-based LLM Inference with Speculative Decoding
von: Yin, Haofei, et al.
Veröffentlicht: (2025)
von: Yin, Haofei, et al.
Veröffentlicht: (2025)
LLM-42: Enabling Determinism in LLM Inference with Verified Speculation
von: Gond, Raja, et al.
Veröffentlicht: (2026)
von: Gond, Raja, et al.
Veröffentlicht: (2026)
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency
von: Li, Ruixiao, et al.
Veröffentlicht: (2025)
von: Li, Ruixiao, et al.
Veröffentlicht: (2025)
Compiler-Assisted Speculative Sampling for Accelerated LLM Inference on Heterogeneous Edge Devices
von: Mesa, Alejandro Ruiz y, et al.
Veröffentlicht: (2026)
von: Mesa, Alejandro Ruiz y, et al.
Veröffentlicht: (2026)
Speculative Streaming: Fast LLM Inference without Auxiliary Models
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2024)
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2024)
SOLARIS: Speculative Offloading of Latent-bAsed Representation for Inference Scaling
von: Liu, Zikun, et al.
Veröffentlicht: (2026)
von: Liu, Zikun, et al.
Veröffentlicht: (2026)
Lever LM: Configuring In-Context Sequence to Lever Large Vision Language Models
von: Yang, Xu, et al.
Veröffentlicht: (2023)
von: Yang, Xu, et al.
Veröffentlicht: (2023)
Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
von: Jeon, Wonseok, et al.
Veröffentlicht: (2024)
von: Jeon, Wonseok, et al.
Veröffentlicht: (2024)
GELATO: Generative Entropy- and Lyapunov-based Adaptive Token Offloading for Device-Edge Speculative LLM Inference
von: Tang, Zengzipeng, et al.
Veröffentlicht: (2026)
von: Tang, Zengzipeng, et al.
Veröffentlicht: (2026)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
Fast Inference via Hierarchical Speculative Decoding
von: Mohri, Clara, et al.
Veröffentlicht: (2025)
von: Mohri, Clara, et al.
Veröffentlicht: (2025)
Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding
von: Park, Jihoon, et al.
Veröffentlicht: (2025)
von: Park, Jihoon, et al.
Veröffentlicht: (2025)
Data Distribution as a Lever for Guiding Optimizers Toward Superior Generalization in LLMs
von: Gangavarapu, Tushaar, et al.
Veröffentlicht: (2026)
von: Gangavarapu, Tushaar, et al.
Veröffentlicht: (2026)
PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation
von: Butler, Branden, et al.
Veröffentlicht: (2024)
von: Butler, Branden, et al.
Veröffentlicht: (2024)
Speculating Experts Accelerates Inference for Mixture-of-Experts
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
Calibrated Speculative Decoding: Frequency-Guided Candidate Selection for Efficient Inference
von: Zhou, Xuwen, et al.
Veröffentlicht: (2026)
von: Zhou, Xuwen, et al.
Veröffentlicht: (2026)
Speculative Speculative Decoding
von: Kumar, Tanishq, et al.
Veröffentlicht: (2026)
von: Kumar, Tanishq, et al.
Veröffentlicht: (2026)
PRISM: Parametrically Refactoring Inference for Speculative Sampling Draft Models
von: Wang, Xuliang, et al.
Veröffentlicht: (2026)
von: Wang, Xuliang, et al.
Veröffentlicht: (2026)
Data as a Lever: A Neighbouring Datasets Perspective on Predictive Multiplicity
von: Ganesh, Prakhar, et al.
Veröffentlicht: (2025)
von: Ganesh, Prakhar, et al.
Veröffentlicht: (2025)
CARD: A Cache-Assisted Parallel Speculative Decoding Framework via Query-and-Correct Paradigm for Accelerating LLM Inference
von: Zhou, Enyu, et al.
Veröffentlicht: (2025)
von: Zhou, Enyu, et al.
Veröffentlicht: (2025)
PowerInfer-2: Fast Large Language Model Inference on a Smartphone
von: Xue, Zhenliang, et al.
Veröffentlicht: (2024)
von: Xue, Zhenliang, et al.
Veröffentlicht: (2024)
Distributed Speculative Inference (DSI): Speculation Parallelism for Provably Faster Lossless Language Model Inference
von: Timor, Nadav, et al.
Veröffentlicht: (2024)
von: Timor, Nadav, et al.
Veröffentlicht: (2024)
Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
The N-Grammys: Accelerating Autoregressive Inference with Learning-Free Batched Speculation
von: Stewart, Lawrence, et al.
Veröffentlicht: (2024)
von: Stewart, Lawrence, et al.
Veröffentlicht: (2024)
FastEagle: Cascaded Drafting for Accelerating Speculative Decoding
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
von: Kong, Linghao, et al.
Veröffentlicht: (2026)
von: Kong, Linghao, et al.
Veröffentlicht: (2026)
Evaluating Federated Learning for Cross-Country Mood Inference from Smartphone Sensing Data
von: Kalpande, Sharmad, et al.
Veröffentlicht: (2026)
von: Kalpande, Sharmad, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DynaKV: Enabling Accurate and Efficient Long-Sequence LLM Decoding on Smartphones
von: Wang, Tuowei, et al.
Veröffentlicht: (2025) -
Neuralink: Fast LLM Inference on Smartphones with Neuron Co-Activation Linking
von: Wang, Tuowei, et al.
Veröffentlicht: (2024) -
Collaborative Speculative Inference for Efficient LLM Inference Serving
von: Gao, Luyao, et al.
Veröffentlicht: (2025) -
Cascade Speculative Drafting for Even Faster LLM Inference
von: Chen, Ziyi, et al.
Veröffentlicht: (2023) -
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2025)