Memory Analysis on the Training Course of DeepSeek Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Ping, Su, Lei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating the Performance of the DeepSeek Model in Confidential Computing Environment
di: Dong, Ben, et al.
Pubblicazione: (2025)
di: Dong, Ben, et al.
Pubblicazione: (2025)
Forecasting GPU Performance for Deep Learning Training and Inference
di: Lee, Seonho, et al.
Pubblicazione: (2024)
di: Lee, Seonho, et al.
Pubblicazione: (2024)
V-Seek: Accelerating LLM Reasoning on Open-hardware Server-class RISC-V Platforms
di: Rodrigo, Javier J. Poveda, et al.
Pubblicazione: (2025)
di: Rodrigo, Javier J. Poveda, et al.
Pubblicazione: (2025)
Greener Deep Reinforcement Learning: Analysis of Energy and Carbon Efficiency Across Atari Benchmarks
di: Gardner, Jason, et al.
Pubblicazione: (2025)
di: Gardner, Jason, et al.
Pubblicazione: (2025)
ZO2: Scalable Zeroth-Order Fine-Tuning for Extremely Large Language Models with Limited GPU Memory
di: Wang, Liangyu, et al.
Pubblicazione: (2025)
di: Wang, Liangyu, et al.
Pubblicazione: (2025)
xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
di: Shi, Jiabo, et al.
Pubblicazione: (2025)
di: Shi, Jiabo, et al.
Pubblicazione: (2025)
Biases in Edge Language Models: Detection, Analysis, and Mitigation
di: Sharma, Vinamra, et al.
Pubblicazione: (2025)
di: Sharma, Vinamra, et al.
Pubblicazione: (2025)
DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
di: Holmes, Connor, et al.
Pubblicazione: (2024)
di: Holmes, Connor, et al.
Pubblicazione: (2024)
oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning Compilation
di: Li, Jianhui, et al.
Pubblicazione: (2023)
di: Li, Jianhui, et al.
Pubblicazione: (2023)
Tabular and Deep Reinforcement Learning for Gittins Index
di: Dhankhar, Harshit, et al.
Pubblicazione: (2024)
di: Dhankhar, Harshit, et al.
Pubblicazione: (2024)
Anatomizing Deep Learning Inference in Web Browsers
di: Wang, Qipeng, et al.
Pubblicazione: (2024)
di: Wang, Qipeng, et al.
Pubblicazione: (2024)
DistZO2: High-Throughput and Memory-Efficient Zeroth-Order Fine-tuning LLMs with Distributed Parallel Computing
di: Wang, Liangyu, et al.
Pubblicazione: (2025)
di: Wang, Liangyu, et al.
Pubblicazione: (2025)
Quantitative Analysis of Performance Drop in DeepSeek Model Quantization
di: Zhao, Enbo, et al.
Pubblicazione: (2025)
di: Zhao, Enbo, et al.
Pubblicazione: (2025)
A High-Throughput Compute-Efficient POMDP Hide-And-Seek-Engine (HASE) for Multi-Agent Operations
di: Flavin, Timothy, et al.
Pubblicazione: (2026)
di: Flavin, Timothy, et al.
Pubblicazione: (2026)
Conformer-Based Speech Recognition On Extreme Edge-Computing Devices
di: Xu, Mingbin, et al.
Pubblicazione: (2023)
di: Xu, Mingbin, et al.
Pubblicazione: (2023)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
di: Shao, Zishan, et al.
Pubblicazione: (2025)
di: Shao, Zishan, et al.
Pubblicazione: (2025)
PreLoRA: Hybrid Pre-training of Vision Transformers with Full Training and Low-Rank Adapters
di: Thapa, Krishu K, et al.
Pubblicazione: (2025)
di: Thapa, Krishu K, et al.
Pubblicazione: (2025)
STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal Planning
di: Huang, Zixiao, et al.
Pubblicazione: (2025)
di: Huang, Zixiao, et al.
Pubblicazione: (2025)
APOLLO: SGD-like Memory, AdamW-level Performance
di: Zhu, Hanqing, et al.
Pubblicazione: (2024)
di: Zhu, Hanqing, et al.
Pubblicazione: (2024)
A Structure-Aware Framework for Learning Device Placements on Computation Graphs
di: Duan, Shukai, et al.
Pubblicazione: (2024)
di: Duan, Shukai, et al.
Pubblicazione: (2024)
P-MOSS: Scheduling Main-Memory Indexes Over NUMA Servers Using Next Token Prediction
di: Rayhan, Yeasir, et al.
Pubblicazione: (2024)
di: Rayhan, Yeasir, et al.
Pubblicazione: (2024)
The Race to Efficiency: A New Perspective on AI Scaling Laws
di: Lu, Chien-Ping
Pubblicazione: (2025)
di: Lu, Chien-Ping
Pubblicazione: (2025)
ProTrain: Efficient LLM Training via Memory-Aware Techniques
di: Yang, Hanmei, et al.
Pubblicazione: (2024)
di: Yang, Hanmei, et al.
Pubblicazione: (2024)
Towards Computational Performance Engineering for Unsupervised Concept Drift Detection -- Complexities, Benchmarking, Performance Analysis
di: Werner, Elias, et al.
Pubblicazione: (2023)
di: Werner, Elias, et al.
Pubblicazione: (2023)
Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
di: You, Bozhi, et al.
Pubblicazione: (2025)
di: You, Bozhi, et al.
Pubblicazione: (2025)
Data Efficacy for Language Model Training
di: Dai, Yalun, et al.
Pubblicazione: (2025)
di: Dai, Yalun, et al.
Pubblicazione: (2025)
Quantifying the Capability Boundary of DeepSeek Models: An Application-Driven Performance Analysis
di: Zhao, Kaikai, et al.
Pubblicazione: (2025)
di: Zhao, Kaikai, et al.
Pubblicazione: (2025)
Can Graph Reordering Speed Up Graph Neural Network Training? An Experimental Study
di: Merkel, Nikolai, et al.
Pubblicazione: (2024)
di: Merkel, Nikolai, et al.
Pubblicazione: (2024)
Benchmarking GPUs on SVBRDF Extractor Model
di: Kandel, Narayan, et al.
Pubblicazione: (2023)
di: Kandel, Narayan, et al.
Pubblicazione: (2023)
Sig2Model: A Boosting-Driven Model for Updatable Learned Indexes
di: Heidari, Alireza, et al.
Pubblicazione: (2025)
di: Heidari, Alireza, et al.
Pubblicazione: (2025)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
di: Kong, Linghao, et al.
Pubblicazione: (2026)
di: Kong, Linghao, et al.
Pubblicazione: (2026)
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
di: Wang, Haoxin, et al.
Pubblicazione: (2025)
di: Wang, Haoxin, et al.
Pubblicazione: (2025)
EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training
di: Yi, Qingao, et al.
Pubblicazione: (2025)
di: Yi, Qingao, et al.
Pubblicazione: (2025)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
di: An, Zihao, et al.
Pubblicazione: (2025)
di: An, Zihao, et al.
Pubblicazione: (2025)
Prefill vs. Decode Bottlenecks: SRAM-Frequency Tradeoffs and the Memory-Bandwidth Ceiling
di: Atmer, Hannah, et al.
Pubblicazione: (2025)
di: Atmer, Hannah, et al.
Pubblicazione: (2025)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
di: Taneja, Maanas, et al.
Pubblicazione: (2026)
di: Taneja, Maanas, et al.
Pubblicazione: (2026)
Research on Low-Latency Inference and Training Efficiency Optimization for Graph Neural Network and Large Language Model-Based Recommendation Systems
di: Zhao, Yushang, et al.
Pubblicazione: (2025)
di: Zhao, Yushang, et al.
Pubblicazione: (2025)
MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
di: Chitty-Venkata, Krishna Teja, et al.
Pubblicazione: (2025)
di: Chitty-Venkata, Krishna Teja, et al.
Pubblicazione: (2025)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
di: Ziller, Thomas, et al.
Pubblicazione: (2026)
di: Ziller, Thomas, et al.
Pubblicazione: (2026)
Deploying Open-Source Large Language Models: A performance Analysis
di: Bendi-Ouis, Yannis, et al.
Pubblicazione: (2024)
di: Bendi-Ouis, Yannis, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Evaluating the Performance of the DeepSeek Model in Confidential Computing Environment
di: Dong, Ben, et al.
Pubblicazione: (2025) -
Forecasting GPU Performance for Deep Learning Training and Inference
di: Lee, Seonho, et al.
Pubblicazione: (2024) -
V-Seek: Accelerating LLM Reasoning on Open-hardware Server-class RISC-V Platforms
di: Rodrigo, Javier J. Poveda, et al.
Pubblicazione: (2025) -
Greener Deep Reinforcement Learning: Analysis of Energy and Carbon Efficiency Across Atari Benchmarks
di: Gardner, Jason, et al.
Pubblicazione: (2025) -
ZO2: Scalable Zeroth-Order Fine-Tuning for Extremely Large Language Models with Limited GPU Memory
di: Wang, Liangyu, et al.
Pubblicazione: (2025)