EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training
Fuente:
arXiv
Salvato in:
| Autori principali: | Yi, Qingao, Duan, Jiaang, Hu, Hanwen, Hua, Qin, Zhao, Haiyan, Qian, Shiyou, Yang, Dingyu, Cao, Jian, Tang, Jinghua, Yu, Yinghao, Liao, Chenzhi, Wang, Kangjin, Zhang, Liping |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Atys: An Efficient Profiling Framework for Identifying Hotspot Functions in Large-scale Cloud Microservices
di: Sun, Jiaqi, et al.
Pubblicazione: (2025)
di: Sun, Jiaqi, et al.
Pubblicazione: (2025)
MOPAR: A Model Partitioning Framework for Deep Learning Inference Services on Serverless Platforms
di: Duan, Jiaang, et al.
Pubblicazione: (2024)
di: Duan, Jiaang, et al.
Pubblicazione: (2024)
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
di: Duan, Jiaang, et al.
Pubblicazione: (2025)
di: Duan, Jiaang, et al.
Pubblicazione: (2025)
WaveTune: Wave-aware Bilinear Modeling for Efficient GPU Kernel Auto-tuning
di: Zhang, Kaixuan, et al.
Pubblicazione: (2026)
di: Zhang, Kaixuan, et al.
Pubblicazione: (2026)
PipeWeave: Synergizing Analytical and Learning Models for Unified GPU Performance Prediction
di: Zhang, Kaixuan, et al.
Pubblicazione: (2026)
di: Zhang, Kaixuan, et al.
Pubblicazione: (2026)
GreedySnake: Accelerating SSD-Offloaded LLM Training with Efficient Scheduling and Optimizer Step Overlapping
di: Yin, Yishu, et al.
Pubblicazione: (2025)
di: Yin, Yishu, et al.
Pubblicazione: (2025)
SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference
di: Shin, Jiho, et al.
Pubblicazione: (2024)
di: Shin, Jiho, et al.
Pubblicazione: (2024)
Waltz: Temperature-Aware Cooperative Compression for High-Performance Compression-Based CSDs
di: Yu, Dingcui, et al.
Pubblicazione: (2025)
di: Yu, Dingcui, et al.
Pubblicazione: (2025)
SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving
di: Zhang, Quqing, et al.
Pubblicazione: (2026)
di: Zhang, Quqing, et al.
Pubblicazione: (2026)
Fast Entropy Decoding for Sparse MVM on GPUs
di: Schätzle, Emil, et al.
Pubblicazione: (2026)
di: Schätzle, Emil, et al.
Pubblicazione: (2026)
SysOM-AI: Continuous Cross-Layer Performance Diagnosis for Production AI Training
di: Zheng, Yusheng, et al.
Pubblicazione: (2026)
di: Zheng, Yusheng, et al.
Pubblicazione: (2026)
Accurate Performance Modeling And Uncertainty Analysis of Lossy Compression in Scientific Applications
di: Liu, Youyuan, et al.
Pubblicazione: (2024)
di: Liu, Youyuan, et al.
Pubblicazione: (2024)
BitLogic: Training Framework for Gradient-Based FPGA-Native Neural Networks
di: Bührer, Simon, et al.
Pubblicazione: (2026)
di: Bührer, Simon, et al.
Pubblicazione: (2026)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
di: Liu, Guangda, et al.
Pubblicazione: (2024)
di: Liu, Guangda, et al.
Pubblicazione: (2024)
GreenLLM: SLO-Aware Dynamic Frequency Scaling for Energy-Efficient LLM Serving
di: Liu, Qunyou, et al.
Pubblicazione: (2025)
di: Liu, Qunyou, et al.
Pubblicazione: (2025)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
di: Karfakis, George, et al.
Pubblicazione: (2025)
di: Karfakis, George, et al.
Pubblicazione: (2025)
Profiling Apple Silicon Performance for ML Training
di: Feng, Dahua, et al.
Pubblicazione: (2025)
di: Feng, Dahua, et al.
Pubblicazione: (2025)
A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving
di: Agullo, Ferran, et al.
Pubblicazione: (2025)
di: Agullo, Ferran, et al.
Pubblicazione: (2025)
Evaluating the Performance of the DeepSeek Model in Confidential Computing Environment
di: Dong, Ben, et al.
Pubblicazione: (2025)
di: Dong, Ben, et al.
Pubblicazione: (2025)
SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
di: Mozaffari, Mohammad, et al.
Pubblicazione: (2024)
di: Mozaffari, Mohammad, et al.
Pubblicazione: (2024)
Statistical Modeling and Uncertainty Estimation of LLM Inference Systems
di: Ray, Kaustabha, et al.
Pubblicazione: (2025)
di: Ray, Kaustabha, et al.
Pubblicazione: (2025)
Characterize LSM-tree Compaction Performance via On-Device LLM Inference
di: Ding, Jiabiao, et al.
Pubblicazione: (2026)
di: Ding, Jiabiao, et al.
Pubblicazione: (2026)
LoPace: A Lossless Optimized Prompt Accurate Compression Engine for Large Language Model Applications
di: Ulla, Aman
Pubblicazione: (2026)
di: Ulla, Aman
Pubblicazione: (2026)
RWKV-edge: Deeply Compressed RWKV for Resource-Constrained Devices
di: Choe, Wonkyo, et al.
Pubblicazione: (2024)
di: Choe, Wonkyo, et al.
Pubblicazione: (2024)
Humas: A Heterogeneity- and Upgrade-aware Microservice Auto-scaling Framework in Large-scale Data Centers
di: Hua, Qin, et al.
Pubblicazione: (2024)
di: Hua, Qin, et al.
Pubblicazione: (2024)
Single 32-bit Sub-Channel DDR5 DIMMs: Architecture, Performance Bounds, and Standardisation
di: Ke, Chih-Hua
Pubblicazione: (2026)
di: Ke, Chih-Hua
Pubblicazione: (2026)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
di: Fu, Zizhuo, et al.
Pubblicazione: (2025)
di: Fu, Zizhuo, et al.
Pubblicazione: (2025)
Affine Frequency Division Multiplexing Over Wideband Doubly-Dispersive Channels With Time-Scaling Effects
di: Li, Xiangxiang, et al.
Pubblicazione: (2025)
di: Li, Xiangxiang, et al.
Pubblicazione: (2025)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
di: Taneja, Maanas, et al.
Pubblicazione: (2026)
di: Taneja, Maanas, et al.
Pubblicazione: (2026)
Toward Greener Matrix Operations by Lossless Compressed Formats
di: Tosoni, Francesco, et al.
Pubblicazione: (2024)
di: Tosoni, Francesco, et al.
Pubblicazione: (2024)
Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage
di: Yuan, Ziqi, et al.
Pubblicazione: (2025)
di: Yuan, Ziqi, et al.
Pubblicazione: (2025)
Training Transformers in Cosine Coefficient Space
di: Bergach, Mohamed Amine
Pubblicazione: (2026)
di: Bergach, Mohamed Amine
Pubblicazione: (2026)
A Model-driven Approach for Continuous Performance Engineering in Microservice-based Systems
di: Cortellessa, Vittorio, et al.
Pubblicazione: (2023)
di: Cortellessa, Vittorio, et al.
Pubblicazione: (2023)
FRSZ2 for In-Register Block Compression Inside GMRES on GPUs
di: Grützmacher, Thomas, et al.
Pubblicazione: (2024)
di: Grützmacher, Thomas, et al.
Pubblicazione: (2024)
LKD-KGC: Domain-Specific KG Construction via LLM-driven Knowledge Dependency Parsing
di: Sun, Jiaqi, et al.
Pubblicazione: (2025)
di: Sun, Jiaqi, et al.
Pubblicazione: (2025)
Dawn of the Dead(line Misses): Impact of Job Dismiss on the Deadline Miss Rate
di: Chen, Jian-Jia, et al.
Pubblicazione: (2024)
di: Chen, Jian-Jia, et al.
Pubblicazione: (2024)
ProTrain: Efficient LLM Training via Memory-Aware Techniques
di: Yang, Hanmei, et al.
Pubblicazione: (2024)
di: Yang, Hanmei, et al.
Pubblicazione: (2024)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
di: Zhang, Li, et al.
Pubblicazione: (2025)
di: Zhang, Li, et al.
Pubblicazione: (2025)
Memory Analysis on the Training Course of DeepSeek Models
di: Zhang, Ping, et al.
Pubblicazione: (2025)
di: Zhang, Ping, et al.
Pubblicazione: (2025)
Forecasting GPU Performance for Deep Learning Training and Inference
di: Lee, Seonho, et al.
Pubblicazione: (2024)
di: Lee, Seonho, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Atys: An Efficient Profiling Framework for Identifying Hotspot Functions in Large-scale Cloud Microservices
di: Sun, Jiaqi, et al.
Pubblicazione: (2025) -
MOPAR: A Model Partitioning Framework for Deep Learning Inference Services on Serverless Platforms
di: Duan, Jiaang, et al.
Pubblicazione: (2024) -
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
di: Duan, Jiaang, et al.
Pubblicazione: (2025) -
WaveTune: Wave-aware Bilinear Modeling for Efficient GPU Kernel Auto-tuning
di: Zhang, Kaixuan, et al.
Pubblicazione: (2026) -
PipeWeave: Synergizing Analytical and Learning Models for Unified GPU Performance Prediction
di: Zhang, Kaixuan, et al.
Pubblicazione: (2026)