LLaMCAT: Optimizing Large Language Model Inference with Cache Arbitration and Throttling
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Zhongchun, Lai, Chengtao, Zhang, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DCO: Dynamic Cache Orchestration for LLM Accelerators through Predictive Management
by: Zhou, Zhongchun, et al.
Published: (2025)
by: Zhou, Zhongchun, et al.
Published: (2025)
Cache Your Prompt When It's Green: Carbon-Aware Caching for Large Language Model Serving
by: Tian, Yuyang, et al.
Published: (2025)
by: Tian, Yuyang, et al.
Published: (2025)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
by: Xu, Weihong, et al.
Published: (2025)
by: Xu, Weihong, et al.
Published: (2025)
Global Optimizations & Lightweight Dynamic Logic for Concurrency
by: Pati, Suchita, et al.
Published: (2024)
by: Pati, Suchita, et al.
Published: (2024)
Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
by: Zheng, Xianzhe, et al.
Published: (2026)
by: Zheng, Xianzhe, et al.
Published: (2026)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
by: li, Fei, et al.
Published: (2026)
by: li, Fei, et al.
Published: (2026)
FengHuang: Next-Generation Memory Orchestration for AI Inferencing
by: Li, Jiamin, et al.
Published: (2025)
by: Li, Jiamin, et al.
Published: (2025)
Pooling Engram Conditional Memory in Large Language Models using CXL
by: Ma, Ruiyang, et al.
Published: (2026)
by: Ma, Ruiyang, et al.
Published: (2026)
LFOC: A Lightweight Fairness-Oriented Cache Clustering Policy for Commodity Multicores
by: García-García, Adrián, et al.
Published: (2024)
by: García-García, Adrián, et al.
Published: (2024)
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
by: Shi, Tianyao, et al.
Published: (2024)
by: Shi, Tianyao, et al.
Published: (2024)
RevaMp3D: Architecting the Processor Core and Cache Hierarchy for Systems with Monolithically-Integrated Logic and Memory
by: Ghiasi, Nika Mansouri, et al.
Published: (2022)
by: Ghiasi, Nika Mansouri, et al.
Published: (2022)
NPU Design for Diffusion Language Model Inference
by: Lou, Binglei, et al.
Published: (2026)
by: Lou, Binglei, et al.
Published: (2026)
Performance Modeling and Workload Analysis of Distributed Large Language Model Training and Inference
by: Kundu, Joyjit, et al.
Published: (2024)
by: Kundu, Joyjit, et al.
Published: (2024)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
by: Chung, Euijun, et al.
Published: (2026)
by: Chung, Euijun, et al.
Published: (2026)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
by: Meng, William, et al.
Published: (2025)
by: Meng, William, et al.
Published: (2025)
Automated Deep Neural Network Inference Partitioning for Distributed Embedded Systems
by: Kreß, Fabian, et al.
Published: (2024)
by: Kreß, Fabian, et al.
Published: (2024)
Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Voxel
by: Liu, Yiqi, et al.
Published: (2026)
by: Liu, Yiqi, et al.
Published: (2026)
HgPCN: A Heterogeneous Architecture for E2E Embedded Point Cloud Inference
by: Gao, Yiming, et al.
Published: (2025)
by: Gao, Yiming, et al.
Published: (2025)
Optimizing Offload Performance in Heterogeneous MPSoCs
by: Colagrande, Luca, et al.
Published: (2024)
by: Colagrande, Luca, et al.
Published: (2024)
Llumnix: Dynamic Scheduling for Large Language Model Serving
by: Sun, Biao, et al.
Published: (2024)
by: Sun, Biao, et al.
Published: (2024)
Optimizing Task Scheduling in Fog Computing with Deadline Awareness
by: Sirjani, Mohammad Sadegh, et al.
Published: (2025)
by: Sirjani, Mohammad Sadegh, et al.
Published: (2025)
Leveraging SIMD for Accelerating Large-number Arithmetic
by: Das, Subhrajit, et al.
Published: (2026)
by: Das, Subhrajit, et al.
Published: (2026)
NetSmith: An Optimization Framework for Machine-Discovered Network Topologies
by: Green, Conor, et al.
Published: (2024)
by: Green, Conor, et al.
Published: (2024)
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
by: Punniyamurthy, Kishore, et al.
Published: (2023)
by: Punniyamurthy, Kishore, et al.
Published: (2023)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
by: Agrawal, Anirudha, et al.
Published: (2024)
by: Agrawal, Anirudha, et al.
Published: (2024)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
by: Zhou, Zhuoshan, et al.
Published: (2026)
by: Zhou, Zhuoshan, et al.
Published: (2026)
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
by: Shi, Man, et al.
Published: (2024)
by: Shi, Man, et al.
Published: (2024)
Optimizing Communication for Latency Sensitive HPC Applications on up to 48 FPGAs Using ACCL
by: Meyer, Marius, et al.
Published: (2024)
by: Meyer, Marius, et al.
Published: (2024)
Switchboard: An Open-Source Framework for Modular Simulation of Large Hardware Systems
by: Herbst, Steven, et al.
Published: (2024)
by: Herbst, Steven, et al.
Published: (2024)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
by: Zhang, Qijun, et al.
Published: (2026)
by: Zhang, Qijun, et al.
Published: (2026)
FLEX: Leveraging FPGA-CPU Synergy for Mixed-Cell-Height Legalization Acceleration
by: Liu, Xingyu, et al.
Published: (2025)
by: Liu, Xingyu, et al.
Published: (2025)
Machine Learning-Driven Intelligent Memory System Design: From On-Chip Caches to Storage
by: Bera, Rahul, et al.
Published: (2026)
by: Bera, Rahul, et al.
Published: (2026)
Taming Offload Overheads in a Massively Parallel Open-Source RISC-V MPSoC: Analysis and Optimization
by: Colagrande, Luca, et al.
Published: (2025)
by: Colagrande, Luca, et al.
Published: (2025)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
by: Colagrande, Luca, et al.
Published: (2026)
by: Colagrande, Luca, et al.
Published: (2026)
Navigating the Landscape of Distributed File Systems: Architectures, Implementations, and Considerations
by: Pan, Xueting, et al.
Published: (2024)
by: Pan, Xueting, et al.
Published: (2024)
Athena: Synergizing Data Prefetching and Off-Chip Prediction via Online Reinforcement Learning
by: Bera, Rahul, et al.
Published: (2026)
by: Bera, Rahul, et al.
Published: (2026)
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
by: Yüzügüler, Ahmet Caner, et al.
Published: (2025)
by: Yüzügüler, Ahmet Caner, et al.
Published: (2025)
FlexStep: Enabling Flexible Error Detection in Multi/Many-core Real-time Systems
by: Wang, Tinglue, et al.
Published: (2025)
by: Wang, Tinglue, et al.
Published: (2025)
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
by: Lin, Bin, et al.
Published: (2024)
by: Lin, Bin, et al.
Published: (2024)
Similar Items
-
DCO: Dynamic Cache Orchestration for LLM Accelerators through Predictive Management
by: Zhou, Zhongchun, et al.
Published: (2025) -
Cache Your Prompt When It's Green: Carbon-Aware Caching for Large Language Model Serving
by: Tian, Yuyang, et al.
Published: (2025) -
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
by: Xu, Weihong, et al.
Published: (2025) -
Global Optimizations & Lightweight Dynamic Logic for Concurrency
by: Pati, Suchita, et al.
Published: (2024) -
Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
by: Zheng, Xianzhe, et al.
Published: (2026)