Inshrinkerator: Compressing Deep Learning Training Checkpoints via Dynamic Quantization
Fuente:
arXiv
Guardado en:
| Autores principales: | Agrawal, Amey, Reddy, Sameer, Bhattamishra, Satwik, Nookala, Venkata Prabhakara Sarath, Vashishth, Vidushi, Rong, Kexin, Tumanov, Alexey |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Hardness of Learning Regular Languages in the Next Symbol Prediction Setting
por: Bhattamishra, Satwik, et al.
Publicado: (2025)
por: Bhattamishra, Satwik, et al.
Publicado: (2025)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
por: Yarlagadda, Srihas, et al.
Publicado: (2025)
por: Yarlagadda, Srihas, et al.
Publicado: (2025)
Provably Learning Attention with Queries
por: Bhattamishra, Satwik, et al.
Publicado: (2026)
por: Bhattamishra, Satwik, et al.
Publicado: (2026)
Separations in the Representational Capabilities of Transformers and Recurrent Architectures
por: Bhattamishra, Satwik, et al.
Publicado: (2024)
por: Bhattamishra, Satwik, et al.
Publicado: (2024)
Discovering Interpretable Algorithms by Decompiling Transformers to RASP
por: Huang, Xinting, et al.
Publicado: (2026)
por: Huang, Xinting, et al.
Publicado: (2026)
Benefits and Limitations of Communication in Multi-Agent Reasoning
por: Rizvi-Martel, Michael, et al.
Publicado: (2025)
por: Rizvi-Martel, Michael, et al.
Publicado: (2025)
On Evaluating Performance of LLM Inference Serving Systems
por: Agrawal, Amey, et al.
Publicado: (2025)
por: Agrawal, Amey, et al.
Publicado: (2025)
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
por: Agrawal, Amey, et al.
Publicado: (2024)
por: Agrawal, Amey, et al.
Publicado: (2024)
Vidur: A Large-Scale Simulation Framework For LLM Inference
por: Agrawal, Amey, et al.
Publicado: (2024)
por: Agrawal, Amey, et al.
Publicado: (2024)
Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems
por: Agrawal, Amey, et al.
Publicado: (2024)
por: Agrawal, Amey, et al.
Publicado: (2024)
DεpS: Delayed ε-Shrinking for Faster Once-For-All Training
por: Annavajjala, Aditya, et al.
Publicado: (2024)
por: Annavajjala, Aditya, et al.
Publicado: (2024)
No Request Left Behind: Tackling Heterogeneity in Long-Context LLM Inference with Medha
por: Agrawal, Amey, et al.
Publicado: (2024)
por: Agrawal, Amey, et al.
Publicado: (2024)
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
por: Behnam, Payman, et al.
Publicado: (2025)
por: Behnam, Payman, et al.
Publicado: (2025)
Beyond Following: Mixing Active Initiative into Computational Creativity
por: Lin, Zhiyu, et al.
Publicado: (2024)
por: Lin, Zhiyu, et al.
Publicado: (2024)
Beyond Prompts: Exploring the Design Space of Mixed-Initiative Co-Creativity Systems
por: Lin, Zhiyu, et al.
Publicado: (2023)
por: Lin, Zhiyu, et al.
Publicado: (2023)
BitSnap: Checkpoint Sparsification and Quantization in LLM Training
por: Peng, Yanxin, et al.
Publicado: (2025)
por: Peng, Yanxin, et al.
Publicado: (2025)
PLUM: Improving Inference Efficiency By Leveraging Repetition-Sparsity Trade-Off
por: Kuhar, Sachit, et al.
Publicado: (2023)
por: Kuhar, Sachit, et al.
Publicado: (2023)
VillainNet: Targeted Poisoning Attacks Against SuperNets Along the Accuracy-Latency Pareto Frontier - Artifact
por: David Oygenblik, et al.
Publicado: (2025)
por: David Oygenblik, et al.
Publicado: (2025)
Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving
por: Agrawal, Amey, et al.
Publicado: (2026)
por: Agrawal, Amey, et al.
Publicado: (2026)
Modelling the rotation dependence of cycle variability in sun-like stars: Answering why only slowly rotating stars produce grand minima
por: Vashishth, Vindya
Publicado: (2024)
por: Vashishth, Vindya
Publicado: (2024)
Hysteresis near the transition of the large-scale dynamo in the presence of the small-scale dynamo
por: Vashishth, Vindya
Publicado: (2024)
por: Vashishth, Vindya
Publicado: (2024)
Exploring the neuroprotective role of ubiquinone by attenuation of oxidative stress and mitochondrial dysfunction in an experimental model of myalgic encephalomyelitis/chronic fatigue stress
por: Anushka Vashishth
Publicado: (2024)
por: Anushka Vashishth
Publicado: (2024)
Exploring Quantization for Efficient Pre-Training of Transformer Language Models
por: Chitsaz, Kamran, et al.
Publicado: (2024)
por: Chitsaz, Kamran, et al.
Publicado: (2024)
A Formal Framework for Understanding Length Generalization in Transformers
por: Huang, Xinting, et al.
Publicado: (2024)
por: Huang, Xinting, et al.
Publicado: (2024)
STIQ: Safeguarding Training and Inferencing of Quantum Neural Networks from Untrusted Cloud
por: Kundu, Satwik, et al.
Publicado: (2024)
por: Kundu, Satwik, et al.
Publicado: (2024)
SuperFedNAS: Cost-Efficient Federated Neural Architecture Search for On-Device Inference
por: Khare, Alind, et al.
Publicado: (2023)
por: Khare, Alind, et al.
Publicado: (2023)
Unsteady Free Convection MHD Flow of an Incompressible Electrically Conducting Viscous Fluid through Porous Medium between Two Vertical Plates
por: G.Prabhakara Rao
Publicado: (2014)
por: G.Prabhakara Rao
Publicado: (2014)
Multichromophoric Macrocycle Nanoparticles for Mitigating Cyanine Limit and Efficient Aqueous Photocatalysis via Sequential Energy Transfer
por: Vidushi Gupta, et al.
Publicado: (2025)
por: Vidushi Gupta, et al.
Publicado: (2025)
DeCAL Tokenwise Compression
por: Panwar, Sameer
Publicado: (2025)
por: Panwar, Sameer
Publicado: (2025)
Conditions for sectoriality and compactness of the resolvent for a non-self-adjoint Sturm--Liouville operator with singular distributional potential
por: Tumanov, Sergey N.
Publicado: (2025)
por: Tumanov, Sergey N.
Publicado: (2025)
On Molchanov's criterion for compactness of the resolvent for a non self-adjoint Sturm-Liouville operator
por: Tumanov, Sergey N.
Publicado: (2024)
por: Tumanov, Sergey N.
Publicado: (2024)
An Efficient Compression of Deep Neural Network Checkpoints Based on Prediction and Context Modeling
por: Kim, Yuriy, et al.
Publicado: (2025)
por: Kim, Yuriy, et al.
Publicado: (2025)
Security Concerns in Quantum Machine Learning as a Service
por: Kundu, Satwik, et al.
Publicado: (2024)
por: Kundu, Satwik, et al.
Publicado: (2024)
Adversarial Data Poisoning Attacks on Quantum Machine Learning in the NISQ Era
por: Kundu, Satwik, et al.
Publicado: (2024)
por: Kundu, Satwik, et al.
Publicado: (2024)
Dry Spell Dynamics Impacting the Productivity of Rainfed Crops Over the Semi‐Arid Regions of South‐East India
por: Santanu Kumar Bal, et al.
Publicado: (2024)
por: Santanu Kumar Bal, et al.
Publicado: (2024)
The Transformer Cookbook
por: Yang, Andy, et al.
Publicado: (2025)
por: Yang, Andy, et al.
Publicado: (2025)
Effects of Epoxy Composition on the Thermal and Network Properties of Crosslinked Thermosets: A Molecular-Dynamics Study
por: Pola, Venkata Rama Manoj, et al.
Publicado: (2025)
por: Pola, Venkata Rama Manoj, et al.
Publicado: (2025)
Deep Probabilistic Unfolding for Quantized Compressive Sensing
por: Qu, Gang, et al.
Publicado: (2026)
por: Qu, Gang, et al.
Publicado: (2026)
Faithfulness Measurable Masked Language Models
por: Madsen, Andreas, et al.
Publicado: (2023)
por: Madsen, Andreas, et al.
Publicado: (2023)
Are self-explanations from Large Language Models faithful?
por: Madsen, Andreas, et al.
Publicado: (2024)
por: Madsen, Andreas, et al.
Publicado: (2024)
Ejemplares similares
-
Hardness of Learning Regular Languages in the Next Symbol Prediction Setting
por: Bhattamishra, Satwik, et al.
Publicado: (2025) -
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
por: Yarlagadda, Srihas, et al.
Publicado: (2025) -
Provably Learning Attention with Queries
por: Bhattamishra, Satwik, et al.
Publicado: (2026) -
Separations in the Representational Capabilities of Transformers and Recurrent Architectures
por: Bhattamishra, Satwik, et al.
Publicado: (2024) -
Discovering Interpretable Algorithms by Decompiling Transformers to RASP
por: Huang, Xinting, et al.
Publicado: (2026)