StepCache: Step-Level Reuse with Lightweight Verification and Selective Patching for LLM Serving
Fuente:
arXiv
Guardado en:
| Autor principal: | Nouri, Azam |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving
por: Yang, Mingyu, et al.
Publicado: (2025)
por: Yang, Mingyu, et al.
Publicado: (2025)
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
por: Li, Xiangchen, et al.
Publicado: (2026)
por: Li, Xiangchen, et al.
Publicado: (2026)
Parameter-Efficient and Personalized Federated Training of Generative Models at the Edge
por: Khan, Kabir, et al.
Publicado: (2025)
por: Khan, Kabir, et al.
Publicado: (2025)
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
por: Li, Xiangchen, et al.
Publicado: (2026)
por: Li, Xiangchen, et al.
Publicado: (2026)
Addressing tokens dynamic generation, propagation, storage and renewal to secure the GlideinWMS pilot based jobs and system
por: Coimbra, Bruno Moreira, et al.
Publicado: (2025)
por: Coimbra, Bruno Moreira, et al.
Publicado: (2025)
Scheduling the Unschedulable: Taming Black-Box LLM Inference at Scale
por: Yuan, Renzhong, et al.
Publicado: (2026)
por: Yuan, Renzhong, et al.
Publicado: (2026)
Spark-LLM-Eval: A Distributed Framework for Statistically Rigorous Large Language Model Evaluation
por: Mitra, Subhadip
Publicado: (2026)
por: Mitra, Subhadip
Publicado: (2026)
SLA Management in Reconfigurable Multi-Agent RAG: A Systems Approach to Question Answering
por: Iannelli, Michael, et al.
Publicado: (2024)
por: Iannelli, Michael, et al.
Publicado: (2024)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
por: Kolluru, Saicharan
Publicado: (2025)
por: Kolluru, Saicharan
Publicado: (2025)
Combining Serverless and High-Performance Computing Paradigms to support ML Data-Intensive Applications
por: Staylor, Mills, et al.
Publicado: (2025)
por: Staylor, Mills, et al.
Publicado: (2025)
Deep RC: A Scalable Data Engineering and Deep Learning Pipeline
por: Sarker, Arup Kumar, et al.
Publicado: (2025)
por: Sarker, Arup Kumar, et al.
Publicado: (2025)
Design and Implementation of an Analysis Pipeline for Heterogeneous Data
por: Sarker, Arup Kumar, et al.
Publicado: (2024)
por: Sarker, Arup Kumar, et al.
Publicado: (2024)
Deploy, Calibrate, Monitor, Heal -- No Human Required: An Autonomous AI SRE Agent for Elasticsearch
por: Mukkolakkal, Muhamed Ramees Cheriya
Publicado: (2026)
por: Mukkolakkal, Muhamed Ramees Cheriya
Publicado: (2026)
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
por: Kamath, Aditya K, et al.
Publicado: (2024)
por: Kamath, Aditya K, et al.
Publicado: (2024)
Knowledge Graphs-Driven Intelligence for Distributed Decision Systems
por: Napoli, Rosario, et al.
Publicado: (2026)
por: Napoli, Rosario, et al.
Publicado: (2026)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
por: Penke, Carolin, et al.
Publicado: (2025)
por: Penke, Carolin, et al.
Publicado: (2025)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
por: Ganjihal, Sanjeev Rao
Publicado: (2026)
por: Ganjihal, Sanjeev Rao
Publicado: (2026)
Kant: An Efficient Unified Scheduling System for Large-Scale AI Clusters
por: Zeng, Lingling, et al.
Publicado: (2025)
por: Zeng, Lingling, et al.
Publicado: (2025)
DPDPU: Data Processing with DPUs
por: Hu, Jiasheng, et al.
Publicado: (2024)
por: Hu, Jiasheng, et al.
Publicado: (2024)
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
por: Georgiou, Athos
Publicado: (2026)
por: Georgiou, Athos
Publicado: (2026)
De-DSI: Decentralised Differentiable Search Index
por: Neague, Petru, et al.
Publicado: (2024)
por: Neague, Petru, et al.
Publicado: (2024)
Towards Message Brokers for Generative AI: Survey, Challenges, and Opportunities
por: Saleh, Alaa, et al.
Publicado: (2023)
por: Saleh, Alaa, et al.
Publicado: (2023)
GraphBit: A Graph-based Agentic Framework for Non-Linear Agent Orchestration
por: Sarker, Yeahia, et al.
Publicado: (2026)
por: Sarker, Yeahia, et al.
Publicado: (2026)
Dodoor: Efficient Randomized Decentralized Scheduling with Load Caching for Heterogeneous Tasks and Clusters
por: Da, Wei, et al.
Publicado: (2025)
por: Da, Wei, et al.
Publicado: (2025)
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended
por: Kamath, Aditya K, et al.
Publicado: (2026)
por: Kamath, Aditya K, et al.
Publicado: (2026)
Serverless GPU Architecture for Enterprise HR Analytics: A Production-Scale BDaaS Implementation
por: Zhang, Guilin, et al.
Publicado: (2025)
por: Zhang, Guilin, et al.
Publicado: (2025)
How Machine Learning-Data Driven Replication Strategies Enhance Fault Tolerance in Large-Scale Distributed Systems
por: Murimi, Almond Kiruthu
Publicado: (2025)
por: Murimi, Almond Kiruthu
Publicado: (2025)
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
por: Topcu, Burak, et al.
Publicado: (2026)
por: Topcu, Burak, et al.
Publicado: (2026)
TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
por: Bian, Zhuohang, et al.
Publicado: (2025)
por: Bian, Zhuohang, et al.
Publicado: (2025)
FCDP: Fully Cached Data Parallel for Communication-Avoiding Large-Scale Training
por: Park, Gyeongseo, et al.
Publicado: (2026)
por: Park, Gyeongseo, et al.
Publicado: (2026)
Towards Building Private LLMs: Exploring Multi-Node Expert Parallelism on Apple Silicon for Mixture-of-Experts Large Language Model
por: Chen, Mu-Chi, et al.
Publicado: (2025)
por: Chen, Mu-Chi, et al.
Publicado: (2025)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
por: Jo, Myeong Jun
Publicado: (2026)
por: Jo, Myeong Jun
Publicado: (2026)
Shipwright: Proving liveness of distributed systems with Byzantine participants
por: Leung, Derek, et al.
Publicado: (2025)
por: Leung, Derek, et al.
Publicado: (2025)
Neural Router: Semantic Content Matching for Agentic AI
por: Lovén, Lauri, et al.
Publicado: (2026)
por: Lovén, Lauri, et al.
Publicado: (2026)
EWSJF: An Adaptive Scheduler with Hybrid Partitioning for Mixed-Workload LLM Inference
por: Sidik, Bronislav, et al.
Publicado: (2026)
por: Sidik, Bronislav, et al.
Publicado: (2026)
AMP4EC: Adaptive Model Partitioning Framework for Efficient Deep Learning Inference in Edge Computing Environments
por: Zhang, Guilin, et al.
Publicado: (2025)
por: Zhang, Guilin, et al.
Publicado: (2025)
Experimentally Evaluating the Resource Efficiency of Big Data Autoscaling
por: Will, Jonathan, et al.
Publicado: (2025)
por: Will, Jonathan, et al.
Publicado: (2025)
Service Discovery-Based Hybrid Network Middleware for Efficient Communication in Distributed Robotic Systems
por: Sang, Shiyao, et al.
Publicado: (2025)
por: Sang, Shiyao, et al.
Publicado: (2025)
Deadline-Aware Joint Task Scheduling and Offloading in Mobile Edge Computing Systems
por: Nguyen, Ngoc Hung, et al.
Publicado: (2025)
por: Nguyen, Ngoc Hung, et al.
Publicado: (2025)
Asynchronous Pipeline Parallelism for Real-Time Multilingual Lip Synchronization in Video Communication Systems
por: Caglar, Eren, et al.
Publicado: (2025)
por: Caglar, Eren, et al.
Publicado: (2025)
Ejemplares similares
-
DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving
por: Yang, Mingyu, et al.
Publicado: (2025) -
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
por: Li, Xiangchen, et al.
Publicado: (2026) -
Parameter-Efficient and Personalized Federated Training of Generative Models at the Edge
por: Khan, Kabir, et al.
Publicado: (2025) -
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
por: Li, Xiangchen, et al.
Publicado: (2026) -
Addressing tokens dynamic generation, propagation, storage and renewal to secure the GlideinWMS pilot based jobs and system
por: Coimbra, Bruno Moreira, et al.
Publicado: (2025)