Gespeichert in:
| Hauptverfasser: | Mesa, Alejandro Ruiz y, Korol, Guilherme, Riesterer, Moritz, de Lima, João Paulo Cardoso, Castrillon, Jeronimo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.08060 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Stochastic Depth Training for Adaptive Inference
von: Korol, Guilherme, et al.
Veröffentlicht: (2025)
von: Korol, Guilherme, et al.
Veröffentlicht: (2025)
Efficient In-Memory Acceleration of Sparse Block Diagonal LLMs
von: de Lima, João Paulo Cardoso, et al.
Veröffentlicht: (2025)
von: de Lima, João Paulo Cardoso, et al.
Veröffentlicht: (2025)
Full-Stack Optimization for CAM-Only DNN Inference
von: de Lima, João Paulo C., et al.
Veröffentlicht: (2024)
von: de Lima, João Paulo C., et al.
Veröffentlicht: (2024)
Efficient LLM Inference over Heterogeneous Edge Networks with Speculative Decoding
von: Zhu, Bingjie, et al.
Veröffentlicht: (2025)
von: Zhu, Bingjie, et al.
Veröffentlicht: (2025)
Count2Multiply: Reliable In-Memory High-Radix Counting
von: de Lima, João Paulo Cardoso, et al.
Veröffentlicht: (2024)
von: de Lima, João Paulo Cardoso, et al.
Veröffentlicht: (2024)
Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
The Landscape of Compute-near-memory and Compute-in-memory: A Research and Commercial Overview
von: Khan, Asif Ali, et al.
Veröffentlicht: (2024)
von: Khan, Asif Ali, et al.
Veröffentlicht: (2024)
MING: An Automated CNN-to-Edge MLIR HLS framework
von: Bi, Jiahong, et al.
Veröffentlicht: (2026)
von: Bi, Jiahong, et al.
Veröffentlicht: (2026)
CINM (Cinnamon): A Compilation Infrastructure for Heterogeneous Compute In-Memory and Compute Near-Memory Paradigms
von: Khan, Asif Ali, et al.
Veröffentlicht: (2022)
von: Khan, Asif Ali, et al.
Veröffentlicht: (2022)
Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
von: Jeon, Wonseok, et al.
Veröffentlicht: (2024)
von: Jeon, Wonseok, et al.
Veröffentlicht: (2024)
Quantize-Sample-and-Verify: LLM Acceleration via Adaptive Edge-Cloud Speculative Decoding
von: Zhang, Guangyi, et al.
Veröffentlicht: (2025)
von: Zhang, Guangyi, et al.
Veröffentlicht: (2025)
WISV: Wireless-Informed Semantic Verification for Distributed Speculative Decoding in Device-Edge LLM Inference
von: Liu, Zixuan, et al.
Veröffentlicht: (2026)
von: Liu, Zixuan, et al.
Veröffentlicht: (2026)
Demonstrating a Future for MLIR-native DSL Compilers on a NumPy-like Example
von: Friebel, Karl F. A., et al.
Veröffentlicht: (2026)
von: Friebel, Karl F. A., et al.
Veröffentlicht: (2026)
DSSD: Efficient Edge-Device LLM Deployment and Collaborative Inference via Distributed Split Speculative Decoding
von: Ning, Jiahong, et al.
Veröffentlicht: (2025)
von: Ning, Jiahong, et al.
Veröffentlicht: (2025)
GELATO: Generative Entropy- and Lyapunov-based Adaptive Token Offloading for Device-Edge Speculative LLM Inference
von: Tang, Zengzipeng, et al.
Veröffentlicht: (2026)
von: Tang, Zengzipeng, et al.
Veröffentlicht: (2026)
Hierarchical Verification of Speculative Beams for Accelerating LLM Inference
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
MATCH: Model-Aware TVM-based Compilation for Heterogeneous Edge Devices
von: Hamdi, Mohamed Amine, et al.
Veröffentlicht: (2024)
von: Hamdi, Mohamed Amine, et al.
Veröffentlicht: (2024)
AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration
von: McDanel, Bradley
Veröffentlicht: (2024)
von: McDanel, Bradley
Veröffentlicht: (2024)
SPIN: Accelerating Large Language Model Inference with Heterogeneous Speculative Models
von: Chen, Fahao, et al.
Veröffentlicht: (2025)
von: Chen, Fahao, et al.
Veröffentlicht: (2025)
E-Mapper: Energy-Efficient Resource Allocation for Traditional Operating Systems on Heterogeneous Processors
von: Smejkal, Till, et al.
Veröffentlicht: (2024)
von: Smejkal, Till, et al.
Veröffentlicht: (2024)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
von: Sun, Mingyu, et al.
Veröffentlicht: (2025)
von: Sun, Mingyu, et al.
Veröffentlicht: (2025)
SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration
von: Xia, Heming, et al.
Veröffentlicht: (2024)
von: Xia, Heming, et al.
Veröffentlicht: (2024)
VitaLLM: A Versatile and Tiny Accelerator for Mixed-Precision LLM Inference on Edge Devices
von: Lin, Zi-Wei, et al.
Veröffentlicht: (2026)
von: Lin, Zi-Wei, et al.
Veröffentlicht: (2026)
All-in-Memory Stochastic Computing using ReRAM
von: de Lima, João Paulo C., et al.
Veröffentlicht: (2025)
von: de Lima, João Paulo C., et al.
Veröffentlicht: (2025)
CHIME: Chiplet-based Heterogeneous Near-Memory Acceleration for Edge Multimodal LLM Inference
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
Designing Efficient LLM Accelerators for Edge Devices
von: Haris, Jude, et al.
Veröffentlicht: (2024)
von: Haris, Jude, et al.
Veröffentlicht: (2024)
PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation
von: Butler, Branden, et al.
Veröffentlicht: (2024)
von: Butler, Branden, et al.
Veröffentlicht: (2024)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
CARD: A Cache-Assisted Parallel Speculative Decoding Framework via Query-and-Correct Paradigm for Accelerating LLM Inference
von: Zhou, Enyu, et al.
Veröffentlicht: (2025)
von: Zhou, Enyu, et al.
Veröffentlicht: (2025)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
AHASD: Asynchronous Heterogeneous Architecture for LLM Adaptive Drafting Speculative Decoding on Mobile Devices
von: Zirui, Ma, et al.
Veröffentlicht: (2026)
von: Zirui, Ma, et al.
Veröffentlicht: (2026)
SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed Transmission
von: Zheng, Ce, et al.
Veröffentlicht: (2026)
von: Zheng, Ce, et al.
Veröffentlicht: (2026)
SpecPipe: Accelerating Pipeline Parallelism-based LLM Inference with Speculative Decoding
von: Yin, Haofei, et al.
Veröffentlicht: (2025)
von: Yin, Haofei, et al.
Veröffentlicht: (2025)
Efficiency Unleashed: Inference Acceleration for LLM-based Recommender Systems with Speculative Decoding
von: Xi, Yunjia, et al.
Veröffentlicht: (2024)
von: Xi, Yunjia, et al.
Veröffentlicht: (2024)
CATS: Cascaded Adaptive Tree Speculation for Memory-Limited LLM Inference Acceleration
von: Han, Yuning, et al.
Veröffentlicht: (2026)
von: Han, Yuning, et al.
Veröffentlicht: (2026)
SDSAT: Accelerating LLM Inference through Speculative Decoding with Semantic Adaptive Tokens
von: Liu, Chengbo, et al.
Veröffentlicht: (2024)
von: Liu, Chengbo, et al.
Veröffentlicht: (2024)
Speculative Decoding for Multi-Sample Inference
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
DiP-SD: Distributed Pipelined Speculative Decoding for Efficient LLM Inference at the Edge
von: Xu, Yaodan, et al.
Veröffentlicht: (2026)
von: Xu, Yaodan, et al.
Veröffentlicht: (2026)
CoVSpec: Efficient Device-Edge Co-Inference for Vision-Language Models via Speculative Decoding
von: Jia, Yuanyuan, et al.
Veröffentlicht: (2026)
von: Jia, Yuanyuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Leveraging Stochastic Depth Training for Adaptive Inference
von: Korol, Guilherme, et al.
Veröffentlicht: (2025) -
Efficient In-Memory Acceleration of Sparse Block Diagonal LLMs
von: de Lima, João Paulo Cardoso, et al.
Veröffentlicht: (2025) -
Full-Stack Optimization for CAM-Only DNN Inference
von: de Lima, João Paulo C., et al.
Veröffentlicht: (2024) -
Efficient LLM Inference over Heterogeneous Edge Networks with Speculative Decoding
von: Zhu, Bingjie, et al.
Veröffentlicht: (2025) -
Count2Multiply: Reliable In-Memory High-Radix Counting
von: de Lima, João Paulo Cardoso, et al.
Veröffentlicht: (2024)