Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
Fuente:
arXiv
Guardado en:
| Autores principales: | Kim, Dowon, Lee, MinJae, Kim, Janghyeon, Kwon, HyuckSung, Jeong, Hyeonggyu, Park, Sang-Soo, Yoon, Minyong, Roh, Si-Dong, Kwon, Yongsuk, So, Jinin, Choi, Jungwook |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Cosmos: A CXL-Based Full In-Memory System for Approximate Nearest Neighbor Search
por: Ko, Seoyoung, et al.
Publicado: (2025)
por: Ko, Seoyoung, et al.
Publicado: (2025)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
por: Yoon, Dongha, et al.
Publicado: (2025)
por: Yoon, Dongha, et al.
Publicado: (2025)
KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction
por: Kim, Jang-Hyun, et al.
Publicado: (2025)
por: Kim, Jang-Hyun, et al.
Publicado: (2025)
InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
por: Kim, Minsoo, et al.
Publicado: (2025)
por: Kim, Minsoo, et al.
Publicado: (2025)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
por: Yang, June Yong, et al.
Publicado: (2024)
por: Yang, June Yong, et al.
Publicado: (2024)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
por: Gouk, Donghyun, et al.
Publicado: (2025)
por: Gouk, Donghyun, et al.
Publicado: (2025)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
por: Jo, Dongwon, et al.
Publicado: (2025)
por: Jo, Dongwon, et al.
Publicado: (2025)
CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving
por: Liu, Dong, et al.
Publicado: (2025)
por: Liu, Dong, et al.
Publicado: (2025)
Biomass‐Derived Optically Clear Adhesives for Foldable Displays
por: Youngjoo Park, et al.
Publicado: (2024)
por: Youngjoo Park, et al.
Publicado: (2024)
KVCompose: Efficient Structured KV Cache Compression with Composite Tokens
por: Akulov, Dmitry, et al.
Publicado: (2025)
por: Akulov, Dmitry, et al.
Publicado: (2025)
SPADE: Sparse Pillar-based 3D Object Detection Accelerator for Autonomous Driving
por: Lee, Minjae, et al.
Publicado: (2023)
por: Lee, Minjae, et al.
Publicado: (2023)
LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents
por: Jeon, Hyesung, et al.
Publicado: (2026)
por: Jeon, Hyesung, et al.
Publicado: (2026)
Shared Disk KV Cache Management for Efficient Multi-Instance Inference in RAG-Powered LLMs
por: Lee, Hyungwoo, et al.
Publicado: (2025)
por: Lee, Hyungwoo, et al.
Publicado: (2025)
SEED: Enhancing Text-to-SQL Performance and Practical Usability Through Automatic Evidence Generation
por: Yun, Janghyeon, et al.
Publicado: (2025)
por: Yun, Janghyeon, et al.
Publicado: (2025)
Formalising CXL Cache Coherence
por: Tan, Chengsong, et al.
Publicado: (2024)
por: Tan, Chengsong, et al.
Publicado: (2024)
A unified framework for classical and quantum uncertainty relations using stochastic representations
por: Kwon, Euijoon, et al.
Publicado: (2024)
por: Kwon, Euijoon, et al.
Publicado: (2024)
P‐88: Low reflection antistatic material design for improving ambient contrast ratio of LCD Panel
por: Chang Eun Kim, et al.
Publicado: (2024)
por: Chang Eun Kim, et al.
Publicado: (2024)
Reduction of Ultraviolet‐ and Heat‐Induced Aging Using Betulin‐Loaded Arginine–Caprylate Self‐Assembly: Randomized Double‐Blind Clinical Trials
por: Koo Chul Kwon, et al.
Publicado: (2026)
por: Koo Chul Kwon, et al.
Publicado: (2026)
Nanovesicles for Sensitive Skin Care Developed via Self‐Assembly of Glutamine Linoleate
por: Koo Chul Kwon, et al.
Publicado: (2025)
por: Koo Chul Kwon, et al.
Publicado: (2025)
FairKV: Balancing Per-Head KV Cache for Fast Multi-GPU Inference
por: Zhao, Bingzhe, et al.
Publicado: (2025)
por: Zhao, Bingzhe, et al.
Publicado: (2025)
Efficacy and safety of adding a fourth oral antidiabetic drug versus metformin dose escalation in patients with type 2 diabetes inadequately controlled on triple oral combination therapy ( EFFORT ): A 24‐week, randomized, open‐label, multicenter trial
por: So Ra Kim, et al.
Publicado: (2026)
por: So Ra Kim, et al.
Publicado: (2026)
Infinite-level Fock spaces, crystal bases, and tensor product of extremal weight modules of type $A_{+\infty}$
por: Kwon, Jae-Hoon, et al.
Publicado: (2025)
por: Kwon, Jae-Hoon, et al.
Publicado: (2025)
STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming
por: Jung, MinJae, et al.
Publicado: (2026)
por: Jung, MinJae, et al.
Publicado: (2026)
Mg Doping Strategies for Ni‐Rich Layered Cathode Active Materials in Li‐Ion Batteries
por: Namho Koo, et al.
Publicado: (2026)
por: Namho Koo, et al.
Publicado: (2026)
Fast Adaptation with Kernel and Gradient based Meta Leaning
por: Park, JuneYoung, et al.
Publicado: (2024)
por: Park, JuneYoung, et al.
Publicado: (2024)
Accurate KV Cache Quantization with Outlier Tokens Tracing
por: Su, Yi, et al.
Publicado: (2025)
por: Su, Yi, et al.
Publicado: (2025)
Population‐Based Study on Epidemiological Trends in Interventions for Congenital Heart Disease in Korea Using Nationwide Big Data
por: Jae Sung Son, et al.
Publicado: (2025)
por: Jae Sung Son, et al.
Publicado: (2025)
ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification
por: He, Yefei, et al.
Publicado: (2024)
por: He, Yefei, et al.
Publicado: (2024)
Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs
por: Bui, Ngoc, et al.
Publicado: (2025)
por: Bui, Ngoc, et al.
Publicado: (2025)
Fractional Exhaled Nitric Oxide for Diagnosing Eosinophilic Chronic Rhinosinusitis: A Systematic Review and Meta‐analysis
por: Jae Yoon Lee, et al.
Publicado: (2026)
por: Jae Yoon Lee, et al.
Publicado: (2026)
Basroparib inhibits YAP ‐driven cancers by stabilizing angiomotin
por: Young‐Ju Kwon, et al.
Publicado: (2026)
por: Young‐Ju Kwon, et al.
Publicado: (2026)
Recent Progress in Photocathode Interface Engineering for Photoelectrochemical CO2 Reduction Reaction to C1 or C2+ Products
por: Jae Hak Kim, et al.
Publicado: (2025)
por: Jae Hak Kim, et al.
Publicado: (2025)
EP01.17: Assessing the feasibility of three‐dimensional auto segmentation‐based standard plane detection for fetal brain malformations: a preliminary study
por: G. Kwon, et al.
Publicado: (2024)
por: G. Kwon, et al.
Publicado: (2024)
Large Language Models Facilitate Vision Reflection in Image Classification
por: An, Guoyuan, et al.
Publicado: (2025)
por: An, Guoyuan, et al.
Publicado: (2025)
Joint probability density with radial, tangential, and perturbative forces
por: Jung, Jae-Won, et al.
Publicado: (2024)
por: Jung, Jae-Won, et al.
Publicado: (2024)
EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments
por: Kim, Minsoo, et al.
Publicado: (2025)
por: Kim, Minsoo, et al.
Publicado: (2025)
Sharp regularity of gradient blow-up solutions in the Camassa-Holm equation
por: Kim, Yunjoo, et al.
Publicado: (2024)
por: Kim, Yunjoo, et al.
Publicado: (2024)
Comparing the effectiveness of individual occupation‐based reminiscence therapy at home and in a dementia care centre on cognitive function in older adults with mild dementia: a pilot randomised controlled trial
por: Ha Yeong Jung, et al.
Publicado: (2024)
por: Ha Yeong Jung, et al.
Publicado: (2024)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
por: Kim, Kihyun, et al.
Publicado: (2025)
por: Kim, Kihyun, et al.
Publicado: (2025)
LagKV: Lag-Relative Information of the KV Cache Tells Which Tokens Are Important
por: Liang, Manlai, et al.
Publicado: (2025)
por: Liang, Manlai, et al.
Publicado: (2025)
Ejemplares similares
-
Cosmos: A CXL-Based Full In-Memory System for Approximate Nearest Neighbor Search
por: Ko, Seoyoung, et al.
Publicado: (2025) -
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
por: Yoon, Dongha, et al.
Publicado: (2025) -
KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction
por: Kim, Jang-Hyun, et al.
Publicado: (2025) -
InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
por: Kim, Minsoo, et al.
Publicado: (2025) -
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
por: Yang, June Yong, et al.
Publicado: (2024)