Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Iliakopoulou, Nikoleta, Stojkovic, Jovan, Alverti, Chloe, Xu, Tianyin, Franke, Hubertus, Torrellas, Josep |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Random Adaptive Cache Placement Policy
by: Ahire, Vrushank, et al.
Published: (2025)
by: Ahire, Vrushank, et al.
Published: (2025)
The Hitchhiker's Guide to Programming and Optimizing Cache Coherent Heterogeneous Systems: CXL, NVLink-C2C, and AMD Infinity Fabric
by: Wang, Zixuan, et al.
Published: (2024)
by: Wang, Zixuan, et al.
Published: (2024)
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
by: Stojkovic, Jovan, et al.
Published: (2024)
by: Stojkovic, Jovan, et al.
Published: (2024)
CPU-Limits kill Performance: Time to rethink Resource Control
by: Shetty, Chirag, et al.
Published: (2025)
by: Shetty, Chirag, et al.
Published: (2025)
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
by: Stojkovic, Jovan, et al.
Published: (2024)
by: Stojkovic, Jovan, et al.
Published: (2024)
Optimizing CPU Cache Utilization in Cloud VMs with Accurate Cache Abstraction
by: Tofigh, Mani, et al.
Published: (2025)
by: Tofigh, Mani, et al.
Published: (2025)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
by: Lei, Jianlong, et al.
Published: (2026)
by: Lei, Jianlong, et al.
Published: (2026)
Ultra Ethernet's Design Principles and Architectural Innovations
by: Hoefler, Torsten, et al.
Published: (2025)
by: Hoefler, Torsten, et al.
Published: (2025)
Accelerator-as-a-Service in Public Clouds: An Intra-Host Traffic Management View for Performance Isolation in the Wild
by: Zhao, Jiechen, et al.
Published: (2024)
by: Zhao, Jiechen, et al.
Published: (2024)
Affinity Tailor: Dynamic Locality-Aware Scheduling at Scale
by: Ng, Jin Xin, et al.
Published: (2026)
by: Ng, Jin Xin, et al.
Published: (2026)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
by: Qararyah, Fareed, et al.
Published: (2024)
by: Qararyah, Fareed, et al.
Published: (2024)
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
by: Chakraborty, Abhinaba, et al.
Published: (2025)
by: Chakraborty, Abhinaba, et al.
Published: (2025)
ARCAS: Adaptive Runtime System for Chiplet-Aware Scheduling
by: Fogli, Alessandro, et al.
Published: (2025)
by: Fogli, Alessandro, et al.
Published: (2025)
Demystifying Serverless Costs on Public Platforms: Bridging Billing, Architecture, and OS Scheduling
by: Lin, Changyuan, et al.
Published: (2025)
by: Lin, Changyuan, et al.
Published: (2025)
GigaAPI for GPU Parallelization
by: Suvarna, M., et al.
Published: (2025)
by: Suvarna, M., et al.
Published: (2025)
Performance Analysis of HPC applications on the Aurora Supercomputer: Exploring the Impact of HBM-Enabled Intel Xeon Max CPUs
by: Ibeid, Huda, et al.
Published: (2025)
by: Ibeid, Huda, et al.
Published: (2025)
UPMEM Unleashed: Software Secrets for Speed
by: Chmielewski, Krystian, et al.
Published: (2025)
by: Chmielewski, Krystian, et al.
Published: (2025)
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
by: Luo, Weile, et al.
Published: (2025)
by: Luo, Weile, et al.
Published: (2025)
Experimental Assessment of Containers Running on Top of Virtual Machines
by: Aqasizade, Hossein, et al.
Published: (2024)
by: Aqasizade, Hossein, et al.
Published: (2024)
Compiler Support for Speculation in Decoupled Access/Execute Architectures
by: Szafarczyk, Robert, et al.
Published: (2025)
by: Szafarczyk, Robert, et al.
Published: (2025)
Exploring and Evaluating Real-world CXL: Use Cases and System Adoption
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Parallelizing a modern GPU simulator
by: Huerta, Rodrigo, et al.
Published: (2025)
by: Huerta, Rodrigo, et al.
Published: (2025)
Can Asymmetric Tile Buffering Be Beneficial?
by: Wang, Chengyue, et al.
Published: (2025)
by: Wang, Chengyue, et al.
Published: (2025)
Exploiting long vectors with a CFD code: a co-design show case
by: Blancafort, Marc, et al.
Published: (2024)
by: Blancafort, Marc, et al.
Published: (2024)
Simopt -- Simulation pass for Speculative Optimisation of FPGA-CAD flow
by: Wadhwa, Eashan, et al.
Published: (2024)
by: Wadhwa, Eashan, et al.
Published: (2024)
Multi-Objective Memory Bandwidth Regulation and Cache Partitioning for Multicore Real-Time Systems
by: Sun, Binqi, et al.
Published: (2025)
by: Sun, Binqi, et al.
Published: (2025)
Mitigating GIL Bottlenecks in Edge AI Systems
by: Mandal, Mridankan, et al.
Published: (2026)
by: Mandal, Mridankan, et al.
Published: (2026)
RAID Organizations for Improved Reliability and Performance: A Not Entirely Unbiased Tutorial (1st revision)
by: Thomasian, Alexander
Published: (2024)
by: Thomasian, Alexander
Published: (2024)
SwitchFS: Asynchronous Metadata Updates for Distributed Filesystems with In-Network Coordination
by: Xu, Jingwei, et al.
Published: (2024)
by: Xu, Jingwei, et al.
Published: (2024)
GPUVM: GPU-driven Unified Virtual Memory
by: Nazaraliyev, Nurlan, et al.
Published: (2024)
by: Nazaraliyev, Nurlan, et al.
Published: (2024)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
by: Vellaisamy, Prabhu, et al.
Published: (2025)
by: Vellaisamy, Prabhu, et al.
Published: (2025)
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
by: Stojkovic, Jovan, et al.
Published: (2025)
by: Stojkovic, Jovan, et al.
Published: (2025)
Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers
by: Renney, Harri, et al.
Published: (2026)
by: Renney, Harri, et al.
Published: (2026)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
by: Fan, Ruibo, et al.
Published: (2026)
by: Fan, Ruibo, et al.
Published: (2026)
Vector-Centric Machine Learning Systems: A Cross-Stack Approach
by: Jiang, Wenqi
Published: (2025)
by: Jiang, Wenqi
Published: (2025)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
by: Stojkovic, Jovan, et al.
Published: (2025)
by: Stojkovic, Jovan, et al.
Published: (2025)
Nexus: Transparent I/O Offloading for High-Density Serverless Computing
by: Park, JooYoung, et al.
Published: (2026)
by: Park, JooYoung, et al.
Published: (2026)
Towards CXL Resilience to CPU Failures
by: Psistakis, Antonis, et al.
Published: (2026)
by: Psistakis, Antonis, et al.
Published: (2026)
SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module Accelerators
by: Odema, Mohanad, et al.
Published: (2024)
by: Odema, Mohanad, et al.
Published: (2024)
PyGim: An Efficient Graph Neural Network Library for Real Processing-In-Memory Architectures
by: Giannoula, Christina, et al.
Published: (2024)
by: Giannoula, Christina, et al.
Published: (2024)
Similar Items
-
Random Adaptive Cache Placement Policy
by: Ahire, Vrushank, et al.
Published: (2025) -
The Hitchhiker's Guide to Programming and Optimizing Cache Coherent Heterogeneous Systems: CXL, NVLink-C2C, and AMD Infinity Fabric
by: Wang, Zixuan, et al.
Published: (2024) -
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
by: Stojkovic, Jovan, et al.
Published: (2024) -
CPU-Limits kill Performance: Time to rethink Resource Control
by: Shetty, Chirag, et al.
Published: (2025) -
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
by: Stojkovic, Jovan, et al.
Published: (2024)