Cache Your Prompt When It's Green: Carbon-Aware Caching for Large Language Model Serving
Fuente:
arXiv
Guardado en:
| Autores principales: | Tian, Yuyang, Sun, Desen, Ding, Yi, Liu, Sihang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
por: Shi, Tianyao, et al.
Publicado: (2024)
por: Shi, Tianyao, et al.
Publicado: (2024)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
por: li, Fei, et al.
Publicado: (2026)
por: li, Fei, et al.
Publicado: (2026)
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
por: Yüzügüler, Ahmet Caner, et al.
Publicado: (2025)
por: Yüzügüler, Ahmet Caner, et al.
Publicado: (2025)
Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
por: Zheng, Xianzhe, et al.
Publicado: (2026)
por: Zheng, Xianzhe, et al.
Publicado: (2026)
LLaMCAT: Optimizing Large Language Model Inference with Cache Arbitration and Throttling
por: Zhou, Zhongchun, et al.
Publicado: (2025)
por: Zhou, Zhongchun, et al.
Publicado: (2025)
LFOC: A Lightweight Fairness-Oriented Cache Clustering Policy for Commodity Multicores
por: García-García, Adrián, et al.
Publicado: (2024)
por: García-García, Adrián, et al.
Publicado: (2024)
RevaMp3D: Architecting the Processor Core and Cache Hierarchy for Systems with Monolithically-Integrated Logic and Memory
por: Ghiasi, Nika Mansouri, et al.
Publicado: (2022)
por: Ghiasi, Nika Mansouri, et al.
Publicado: (2022)
Llumnix: Dynamic Scheduling for Large Language Model Serving
por: Sun, Biao, et al.
Publicado: (2024)
por: Sun, Biao, et al.
Publicado: (2024)
PiKV: KV Cache Management System for Mixture of Experts
por: Liu, Dong, et al.
Publicado: (2025)
por: Liu, Dong, et al.
Publicado: (2025)
Pooling Engram Conditional Memory in Large Language Models using CXL
por: Ma, Ruiyang, et al.
Publicado: (2026)
por: Ma, Ruiyang, et al.
Publicado: (2026)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
por: Liu, Lian, et al.
Publicado: (2026)
por: Liu, Lian, et al.
Publicado: (2026)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
por: Meng, William, et al.
Publicado: (2025)
por: Meng, William, et al.
Publicado: (2025)
DCO: Dynamic Cache Orchestration for LLM Accelerators through Predictive Management
por: Zhou, Zhongchun, et al.
Publicado: (2025)
por: Zhou, Zhongchun, et al.
Publicado: (2025)
Random Adaptive Cache Placement Policy
por: Ahire, Vrushank, et al.
Publicado: (2025)
por: Ahire, Vrushank, et al.
Publicado: (2025)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
por: Yu, Yanpeng, et al.
Publicado: (2025)
por: Yu, Yanpeng, et al.
Publicado: (2025)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
por: Xu, Weihong, et al.
Publicado: (2025)
por: Xu, Weihong, et al.
Publicado: (2025)
When Servers Meet Species: A Fab-to-Grave Lens on Computing's Biodiversity Impact
por: Shi, Tianyao, et al.
Publicado: (2025)
por: Shi, Tianyao, et al.
Publicado: (2025)
Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications
por: Li, Jiaxi, et al.
Publicado: (2025)
por: Li, Jiaxi, et al.
Publicado: (2025)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
por: Zhang, Chen, et al.
Publicado: (2026)
por: Zhang, Chen, et al.
Publicado: (2026)
Serving Large Language Models on Huawei CloudMatrix384
por: Zuo, Pengfei, et al.
Publicado: (2025)
por: Zuo, Pengfei, et al.
Publicado: (2025)
Optimizing Task Scheduling in Fog Computing with Deadline Awareness
por: Sirjani, Mohammad Sadegh, et al.
Publicado: (2025)
por: Sirjani, Mohammad Sadegh, et al.
Publicado: (2025)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
por: Adnan, Muhammad, et al.
Publicado: (2024)
por: Adnan, Muhammad, et al.
Publicado: (2024)
Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving
por: Ding, Jianru, et al.
Publicado: (2026)
por: Ding, Jianru, et al.
Publicado: (2026)
Enabling Time-Aware Priority Traffic Management over Distributed FPGA Nodes
por: Scionti, Alberto, et al.
Publicado: (2025)
por: Scionti, Alberto, et al.
Publicado: (2025)
Topology-Aware Virtualization over Inter-Core Connected Neural Processing Units
por: Feng, Dahu, et al.
Publicado: (2025)
por: Feng, Dahu, et al.
Publicado: (2025)
Leveraging SIMD for Accelerating Large-number Arithmetic
por: Das, Subhrajit, et al.
Publicado: (2026)
por: Das, Subhrajit, et al.
Publicado: (2026)
Sequence-Aware Split Heuristic to Mitigate SM Underutilization in FlashAttention-3 Low-Head-Count Decoding
por: Font, Martí Llopart, et al.
Publicado: (2026)
por: Font, Martí Llopart, et al.
Publicado: (2026)
Switchboard: An Open-Source Framework for Modular Simulation of Large Hardware Systems
por: Herbst, Steven, et al.
Publicado: (2024)
por: Herbst, Steven, et al.
Publicado: (2024)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
por: Colagrande, Luca, et al.
Publicado: (2026)
por: Colagrande, Luca, et al.
Publicado: (2026)
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
por: Zhang, Yichao, et al.
Publicado: (2026)
por: Zhang, Yichao, et al.
Publicado: (2026)
Torrent: A Distributed DMA for Efficient and Flexible Point-to-Multipoint Data Movement
por: Deng, Yunhao, et al.
Publicado: (2025)
por: Deng, Yunhao, et al.
Publicado: (2025)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
por: Kong, Fanchen, et al.
Publicado: (2025)
por: Kong, Fanchen, et al.
Publicado: (2025)
EDEA: Efficient Dual-Engine Accelerator for Depthwise Separable Convolution with Direct Data Transfer
por: Chen, Yi, et al.
Publicado: (2025)
por: Chen, Yi, et al.
Publicado: (2025)
RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
por: Chen, Yanru, et al.
Publicado: (2025)
por: Chen, Yanru, et al.
Publicado: (2025)
Multi-Objective Memory Bandwidth Regulation and Cache Partitioning for Multicore Real-Time Systems
por: Sun, Binqi, et al.
Publicado: (2025)
por: Sun, Binqi, et al.
Publicado: (2025)
BlockAMC: Scalable In-Memory Analog Matrix Computing for Solving Linear Systems
por: Pan, Lunshuai, et al.
Publicado: (2024)
por: Pan, Lunshuai, et al.
Publicado: (2024)
CCSS: Hardware-Accelerated RTL Simulation with Fast Combinational Logic Computing and Sequential Logic Synchronization
por: Feng, Weigang, et al.
Publicado: (2025)
por: Feng, Weigang, et al.
Publicado: (2025)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
por: Lei, Jianlong, et al.
Publicado: (2026)
por: Lei, Jianlong, et al.
Publicado: (2026)
COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
por: Negi, Shubham, et al.
Publicado: (2025)
por: Negi, Shubham, et al.
Publicado: (2025)
A Heterogeneous Chiplet Architecture for Accelerating End-to-End Transformer Models
por: Sharma, Harsh, et al.
Publicado: (2023)
por: Sharma, Harsh, et al.
Publicado: (2023)
Ejemplares similares
-
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
por: Shi, Tianyao, et al.
Publicado: (2024) -
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
por: li, Fei, et al.
Publicado: (2026) -
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
por: Yüzügüler, Ahmet Caner, et al.
Publicado: (2025) -
Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
por: Zheng, Xianzhe, et al.
Publicado: (2026) -
LLaMCAT: Optimizing Large Language Model Inference with Cache Arbitration and Throttling
por: Zhou, Zhongchun, et al.
Publicado: (2025)