Saved in:
| Main Authors: | Luo, Shutian, Sadiq, Ali Zafar, Yang, Rui, Zhang, Mingye, Shen, Haiying, Wang, Wei, Cheng, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.19481 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Hitchhiker's Guide to Programming and Optimizing Cache Coherent Heterogeneous Systems: CXL, NVLink-C2C, and AMD Infinity Fabric
by: Wang, Zixuan, et al.
Published: (2024)
by: Wang, Zixuan, et al.
Published: (2024)
Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
by: Chen, Yukang, et al.
Published: (2025)
by: Chen, Yukang, et al.
Published: (2025)
Confidential Serverless Computing
by: Sabanic, Patrick, et al.
Published: (2025)
by: Sabanic, Patrick, et al.
Published: (2025)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
Squeezy: Rapid VM Memory Reclamation for Serverless Functions
by: Nikolos, Orestis Lagkas, et al.
Published: (2024)
by: Nikolos, Orestis Lagkas, et al.
Published: (2024)
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
by: Qiu, Shi, et al.
Published: (2026)
by: Qiu, Shi, et al.
Published: (2026)
Oneiros: KV Cache Optimization through Parameter Remapping for Multi-tenant LLM Serving
by: Li, Ruihao, et al.
Published: (2025)
by: Li, Ruihao, et al.
Published: (2025)
Taiji: A DPU Memory Elasticity Solution for In-production Cloud Environments
by: Zheng, Hao, et al.
Published: (2025)
by: Zheng, Hao, et al.
Published: (2025)
Dirigent: Lightweight Serverless Orchestration
by: Cvetković, Lazar, et al.
Published: (2024)
by: Cvetković, Lazar, et al.
Published: (2024)
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
by: Wu, Yinpeng, et al.
Published: (2026)
by: Wu, Yinpeng, et al.
Published: (2026)
Plug. Play. Persist. Inside a Ready-to-Go Havoc C2 Infrastructure
by: Di Santo, Alessio
Published: (2025)
by: Di Santo, Alessio
Published: (2025)
Optimizing Task Scheduling in Heterogeneous Computing Environments: A Comparative Analysis of CPU, GPU, and ASIC Platforms Using E2C Simulator
by: Mohammadjafari, Ali, et al.
Published: (2024)
by: Mohammadjafari, Ali, et al.
Published: (2024)
Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
by: Xu, Yuhang, et al.
Published: (2025)
by: Xu, Yuhang, et al.
Published: (2025)
Embedded Rust or C Firmware? Lessons from an Industrial Microcontroller Use Case with Ariel OS
by: Thapa, Bipin, et al.
Published: (2026)
by: Thapa, Bipin, et al.
Published: (2026)
Diagonal comparison of ample C*-diagonals
by: Kopsacheilis, Grigoris, et al.
Published: (2024)
by: Kopsacheilis, Grigoris, et al.
Published: (2024)
Taming Serverless Cold Starts Through OS Co-Design
by: Holmes, Ben, et al.
Published: (2025)
by: Holmes, Ben, et al.
Published: (2025)
Evaluating Serverless Machine Learning Performance on Google Cloud Run
by: Khatiwada, Prerana, et al.
Published: (2024)
by: Khatiwada, Prerana, et al.
Published: (2024)
RTP-LLM: High-Performance Alibaba LLM Inference Engine
by: Tan, Boyu, et al.
Published: (2026)
by: Tan, Boyu, et al.
Published: (2026)
SSV: Sparse Speculative Verification for Efficient LLM Inference
by: Wang, Zhibin, et al.
Published: (2026)
by: Wang, Zhibin, et al.
Published: (2026)
Invariant $C^*$-subalgebras of the reduced group $C^*$-algebra
by: Amrutam, Tattwamasi, et al.
Published: (2025)
by: Amrutam, Tattwamasi, et al.
Published: (2025)
Nexus: Transparent I/O Offloading for High-Density Serverless Computing
by: Park, JooYoung, et al.
Published: (2026)
by: Park, JooYoung, et al.
Published: (2026)
Leveraging OS-Level Primitives for Robotic Action Management
by: Zheng, Wenxin, et al.
Published: (2025)
by: Zheng, Wenxin, et al.
Published: (2025)
HACache: Leveraging Read Performance with Cache in a Heterogeneous Array
by: Liu, Jialin, et al.
Published: (2026)
by: Liu, Jialin, et al.
Published: (2026)
On strong shift equivalence for row-finite graphs and C*-algebras
by: Brix, Kevin Aguyar, et al.
Published: (2022)
by: Brix, Kevin Aguyar, et al.
Published: (2022)
Refined moves for structure-preserving isomorphism of graph C*-algebras
by: Eilers, Søren, et al.
Published: (2019)
by: Eilers, Søren, et al.
Published: (2019)
The relative radius of comparison of the crossed product of a non-unital C*-algebra by a finite group
by: Asadi-Vasfi, M. Ali, et al.
Published: (2025)
by: Asadi-Vasfi, M. Ali, et al.
Published: (2025)
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
TrEnv-X: Transparently Share Serverless Execution Environments Across Different Functions and Nodes
by: Huang, Jialiang, et al.
Published: (2025)
by: Huang, Jialiang, et al.
Published: (2025)
Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
by: Chu, Kexin, et al.
Published: (2025)
by: Chu, Kexin, et al.
Published: (2025)
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
by: Prabhu, Ramya, et al.
Published: (2024)
by: Prabhu, Ramya, et al.
Published: (2024)
Demystifying Serverless Costs on Public Platforms: Bridging Billing, Architecture, and OS Scheduling
by: Lin, Changyuan, et al.
Published: (2025)
by: Lin, Changyuan, et al.
Published: (2025)
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
by: Song, Yixin, et al.
Published: (2023)
by: Song, Yixin, et al.
Published: (2023)
Imaginary Machines: A Serverless Model for Cloud Applications
by: Wawrzoniak, Michael, et al.
Published: (2024)
by: Wawrzoniak, Michael, et al.
Published: (2024)
Nanvix: A Multikernel OS Design for High-Density Serverless Deployments
by: Segarra, Carlos, et al.
Published: (2026)
by: Segarra, Carlos, et al.
Published: (2026)
Selfless reduced $C^{*}$-algebras of linear groups
by: Vigdorovich, Itamar
Published: (2026)
by: Vigdorovich, Itamar
Published: (2026)
A characterization of $C^*$-simplicity of countable groups via Poisson boundaries
by: Alpeev, Andrei
Published: (2025)
by: Alpeev, Andrei
Published: (2025)
LLM as a System Service on Mobile Devices
by: Yin, Wangsong, et al.
Published: (2024)
by: Yin, Wangsong, et al.
Published: (2024)
Stationary boundaries on the space of amenable subgroups and C*-simplicity
by: Cascioli, Anna, et al.
Published: (2026)
by: Cascioli, Anna, et al.
Published: (2026)
The ideal intersection property for essential groupoid C*-algebras
by: Kennedy, Matthew, et al.
Published: (2021)
by: Kennedy, Matthew, et al.
Published: (2021)
RUISA Operational Ecosystem Architecture
by: AL Mohtar, Mouayad
Published: (2026)
by: AL Mohtar, Mouayad
Published: (2026)
Similar Items
-
The Hitchhiker's Guide to Programming and Optimizing Cache Coherent Heterogeneous Systems: CXL, NVLink-C2C, and AMD Infinity Fabric
by: Wang, Zixuan, et al.
Published: (2024) -
Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
by: Chen, Yukang, et al.
Published: (2025) -
Confidential Serverless Computing
by: Sabanic, Patrick, et al.
Published: (2025) -
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025) -
Squeezy: Rapid VM Memory Reclamation for Serverless Functions
by: Nikolos, Orestis Lagkas, et al.
Published: (2024)