ESS: An Offload-Centric Latent-Cache Management Architecture for DeepSeek-V3.2-Exp
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Xinhang, Zhang, Chao, He, Jiahuan, Liu, Wei, Zhang, Jianming, Zhou, Wenlong, Li, Xiao, Zeng, Pai, Li, Shiyong, Qian, Yuanpan, Li, Dong, Li, Zhaogeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Automatic BLAS Offloading on Unified Memory Architecture: A Study on NVIDIA Grace-Hopper
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
Revealing the Challenges of Attention-FFN Disaggregation for Modern MoE Models and Hardware Systems
von: Liu, Guowei, et al.
Veröffentlicht: (2026)
von: Liu, Guowei, et al.
Veröffentlicht: (2026)
Adaptive K-PackCache: Cost-Centric Data Caching in Cloud
von: Sarkar, Suvarthi, et al.
Veröffentlicht: (2025)
von: Sarkar, Suvarthi, et al.
Veröffentlicht: (2025)
KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
von: Zhang, Huawei, et al.
Veröffentlicht: (2025)
von: Zhang, Huawei, et al.
Veröffentlicht: (2025)
SPPO:Efficient Long-sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
LLM & HPC:Benchmarking DeepSeek's Performance in High-Performance Computing Tasks
von: Nader, Noujoud, et al.
Veröffentlicht: (2025)
von: Nader, Noujoud, et al.
Veröffentlicht: (2025)
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
von: Zhao, Chenggang, et al.
Veröffentlicht: (2025)
von: Zhao, Chenggang, et al.
Veröffentlicht: (2025)
Mao: Machine learning approach for NUMA optimization in Warehouse Scale Computers
von: Liu, Yueji, et al.
Veröffentlicht: (2024)
von: Liu, Yueji, et al.
Veröffentlicht: (2024)
FedCache: A Knowledge Cache-driven Federated Learning Architecture for Personalized Edge Intelligence
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2023)
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2023)
A Survey of Computation Offloading with Task Types
von: Zhang, Siqi, et al.
Veröffentlicht: (2023)
von: Zhang, Siqi, et al.
Veröffentlicht: (2023)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
von: Ma, Chenxiang, et al.
Veröffentlicht: (2025)
von: Ma, Chenxiang, et al.
Veröffentlicht: (2025)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
von: Zhang, Yuning, et al.
Veröffentlicht: (2025)
von: Zhang, Yuning, et al.
Veröffentlicht: (2025)
Adaptive Cache Management for Complex Storage Systems Using CNN-LSTM-Based Spatiotemporal Prediction
von: Wang, Xiaoye, et al.
Veröffentlicht: (2024)
von: Wang, Xiaoye, et al.
Veröffentlicht: (2024)
PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
von: Liu, Kaiwei, et al.
Veröffentlicht: (2025)
von: Liu, Kaiwei, et al.
Veröffentlicht: (2025)
Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via Semantic-Aware Knowledge Caching
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
Communication-Efficient Serving for Video Diffusion Models with Latent Parallelism
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
von: Wang, Wenfeng, et al.
Veröffentlicht: (2026)
von: Wang, Wenfeng, et al.
Veröffentlicht: (2026)
Proposal of Automatic Offloading Method in Mixed Offloading Destination Environment
von: Yamato, Yoji
Veröffentlicht: (2020)
von: Yamato, Yoji
Veröffentlicht: (2020)
FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework
von: Zhu, Jianian, et al.
Veröffentlicht: (2025)
von: Zhu, Jianian, et al.
Veröffentlicht: (2025)
TURNIP: A "Nondeterministic" GPU Runtime with CPU RAM Offload
von: Ding, Zhimin, et al.
Veröffentlicht: (2024)
von: Ding, Zhimin, et al.
Veröffentlicht: (2024)
A Pilot Study on Tunable Precision Emulation via Automatic BLAS Offloading
von: Liu, Hang, et al.
Veröffentlicht: (2025)
von: Liu, Hang, et al.
Veröffentlicht: (2025)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
von: Wu, Yu, et al.
Veröffentlicht: (2025)
von: Wu, Yu, et al.
Veröffentlicht: (2025)
Distributed Massive MIMO-Aided Task Offloading in Satellite-Terrestrial Integrated Multi-Tier VEC Networks
von: Liu, Yixin, et al.
Veröffentlicht: (2024)
von: Liu, Yixin, et al.
Veröffentlicht: (2024)
Egret: Reinforcement Mechanism for Sequential Computation Offloading in Edge Computing
von: Peng, Haosong, et al.
Veröffentlicht: (2024)
von: Peng, Haosong, et al.
Veröffentlicht: (2024)
To Offload or Not To Offload: Model-driven Comparison of Edge-native and On-device Processing In the Era of Accelerators
von: Ng, Nathan, et al.
Veröffentlicht: (2025)
von: Ng, Nathan, et al.
Veröffentlicht: (2025)
MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
von: Yang, Zheming, et al.
Veröffentlicht: (2026)
von: Yang, Zheming, et al.
Veröffentlicht: (2026)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
von: Kim, Kihyun, et al.
Veröffentlicht: (2025)
von: Kim, Kihyun, et al.
Veröffentlicht: (2025)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
von: Wang, Wenfeng, et al.
Veröffentlicht: (2025)
von: Wang, Wenfeng, et al.
Veröffentlicht: (2025)
CacheFL: Privacy-Preserving and Efficient Federated Cache Model Fine-Tuning for Vision-Language Models
von: Yi, Mengjun, et al.
Veröffentlicht: (2025)
von: Yi, Mengjun, et al.
Veröffentlicht: (2025)
Towards Fast Setup and High Throughput of GPU Serverless Computing
von: Zhao, Han, et al.
Veröffentlicht: (2024)
von: Zhao, Han, et al.
Veröffentlicht: (2024)
Plug & Offload: Transparently Offloading TCP Stack onto Off-path SmartNIC with PnO-TCP
von: Nan, Hailong, et al.
Veröffentlicht: (2025)
von: Nan, Hailong, et al.
Veröffentlicht: (2025)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
von: Qianli, Liu, et al.
Veröffentlicht: (2025)
von: Qianli, Liu, et al.
Veröffentlicht: (2025)
OffloadFS: Leveraging Disaggregated Storage for Computation Offloading
von: Moon, Sungho, et al.
Veröffentlicht: (2026)
von: Moon, Sungho, et al.
Veröffentlicht: (2026)
DWM-RO: Decentralized World Models with Reasoning Offloading for SWIPT-enabled Satellite-Terrestrial HetNets
von: Liu, Guangyuan, et al.
Veröffentlicht: (2025)
von: Liu, Guangyuan, et al.
Veröffentlicht: (2025)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
Caching Aided Multi-Tenant Serverless Computing
von: Qiao, Chu, et al.
Veröffentlicht: (2024)
von: Qiao, Chu, et al.
Veröffentlicht: (2024)
Self-Evolving Distributed Memory Architecture for Scalable AI Systems
von: Li, Zixuan, et al.
Veröffentlicht: (2026)
von: Li, Zixuan, et al.
Veröffentlicht: (2026)
Fine-Grained Vectorized Merge Sorting on RISC-V: From Register to Cache
von: Zhang, Jin, et al.
Veröffentlicht: (2024)
von: Zhang, Jin, et al.
Veröffentlicht: (2024)
Orchestrating Joint Offloading and Scheduling for Low-Latency Edge SLAM
von: Zhang, Yao, et al.
Veröffentlicht: (2025)
von: Zhang, Yao, et al.
Veröffentlicht: (2025)
HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference
von: Tang, Peng, et al.
Veröffentlicht: (2024)
von: Tang, Peng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Automatic BLAS Offloading on Unified Memory Architecture: A Study on NVIDIA Grace-Hopper
von: Li, Junjie, et al.
Veröffentlicht: (2024) -
Revealing the Challenges of Attention-FFN Disaggregation for Modern MoE Models and Hardware Systems
von: Liu, Guowei, et al.
Veröffentlicht: (2026) -
Adaptive K-PackCache: Cost-Centric Data Caching in Cloud
von: Sarkar, Suvarthi, et al.
Veröffentlicht: (2025) -
KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
von: Zhang, Huawei, et al.
Veröffentlicht: (2025) -
SPPO:Efficient Long-sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)