ExCP: Extreme LLM Checkpoint Compression via Weight-Momentum Joint Shrinking
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Wenshuo, Chen, Xinghao, Shu, Han, Tang, Yehui, Wang, Yunhe |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SlimLLM: Accurate Structured Pruning for Large Language Models
di: Guo, Jialong, et al.
Pubblicazione: (2025)
di: Guo, Jialong, et al.
Pubblicazione: (2025)
PocketLLM: Ultimate Compression of Large Language Models via Meta Networks
di: Tian, Ye, et al.
Pubblicazione: (2025)
di: Tian, Ye, et al.
Pubblicazione: (2025)
LLM Data Selection and Utilization via Dynamic Bi-level Optimization
di: Yu, Yang, et al.
Pubblicazione: (2025)
di: Yu, Yang, et al.
Pubblicazione: (2025)
No Time to Waste: Squeeze Time into Channel for Mobile Video Understanding
di: Zhai, Yingjie, et al.
Pubblicazione: (2024)
di: Zhai, Yingjie, et al.
Pubblicazione: (2024)
A Survey on Transformer Compression
di: Tang, Yehui, et al.
Pubblicazione: (2024)
di: Tang, Yehui, et al.
Pubblicazione: (2024)
Kangaroo: Lossless Self-Speculative Decoding via Double Early Exiting
di: Liu, Fangcheng, et al.
Pubblicazione: (2024)
di: Liu, Fangcheng, et al.
Pubblicazione: (2024)
Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression
di: Tang, Yuntian, et al.
Pubblicazione: (2026)
di: Tang, Yuntian, et al.
Pubblicazione: (2026)
TinySAM: Pushing the Envelope for Efficient Segment Anything Model
di: Shu, Han, et al.
Pubblicazione: (2023)
di: Shu, Han, et al.
Pubblicazione: (2023)
Mixture of Lookup Experts
di: Jie, Shibo, et al.
Pubblicazione: (2025)
di: Jie, Shibo, et al.
Pubblicazione: (2025)
DenseMamba: State Space Models with Dense Hidden Connection for Efficient Large Language Models
di: He, Wei, et al.
Pubblicazione: (2024)
di: He, Wei, et al.
Pubblicazione: (2024)
Parameter-Efficient Checkpoint Merging via Metrics-Weighted Averaging
di: Yu, Shi Jie, et al.
Pubblicazione: (2025)
di: Yu, Shi Jie, et al.
Pubblicazione: (2025)
FedLWS: Federated Learning with Adaptive Layer-wise Weight Shrinking
di: Shi, Changlong, et al.
Pubblicazione: (2025)
di: Shi, Changlong, et al.
Pubblicazione: (2025)
Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing
di: Rang, Miao, et al.
Pubblicazione: (2026)
di: Rang, Miao, et al.
Pubblicazione: (2026)
SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch Normalization
di: Guo, Jialong, et al.
Pubblicazione: (2024)
di: Guo, Jialong, et al.
Pubblicazione: (2024)
Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning
di: Jie, Shibo, et al.
Pubblicazione: (2024)
di: Jie, Shibo, et al.
Pubblicazione: (2024)
Text2Weight: Bridging Natural Language and Neural Network Weight Spaces
di: Tian, Bowen, et al.
Pubblicazione: (2025)
di: Tian, Bowen, et al.
Pubblicazione: (2025)
Inshrinkerator: Compressing Deep Learning Training Checkpoints via Dynamic Quantization
di: Agrawal, Amey, et al.
Pubblicazione: (2023)
di: Agrawal, Amey, et al.
Pubblicazione: (2023)
ROOT: Robust Orthogonalized Optimizer for Neural Network Training
di: He, Wei, et al.
Pubblicazione: (2025)
di: He, Wei, et al.
Pubblicazione: (2025)
An Empirical Study of World Model Quantization
di: Fu, Zhongqian, et al.
Pubblicazione: (2026)
di: Fu, Zhongqian, et al.
Pubblicazione: (2026)
BitSnap: Checkpoint Sparsification and Quantization in LLM Training
di: Peng, Yanxin, et al.
Pubblicazione: (2025)
di: Peng, Yanxin, et al.
Pubblicazione: (2025)
DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
di: Yu, Xiaoming, et al.
Pubblicazione: (2026)
di: Yu, Xiaoming, et al.
Pubblicazione: (2026)
Knowledge Distillation Must Account for What It Loses
di: Wang, Wenshuo
Pubblicazione: (2026)
di: Wang, Wenshuo
Pubblicazione: (2026)
Entropy trajectory shape predicts LLM reasoning reliability: A diagnostic study of uncertainty dynamics in chain-of-thought
di: Zhao, Xinghao
Pubblicazione: (2026)
di: Zhao, Xinghao
Pubblicazione: (2026)
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression
di: Malinovskii, Vladimir, et al.
Pubblicazione: (2024)
di: Malinovskii, Vladimir, et al.
Pubblicazione: (2024)
Global Momentum Compression for Sparse Communication in Distributed Learning
di: Shi, Chang-Wei, et al.
Pubblicazione: (2019)
di: Shi, Chang-Wei, et al.
Pubblicazione: (2019)
Accelerating Byzantine-Robust Distributed Learning with Compressed Communication via Double Momentum and Variance Reduction
di: Li, Yanghao, et al.
Pubblicazione: (2026)
di: Li, Yanghao, et al.
Pubblicazione: (2026)
LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
di: Yang, Chenghao, et al.
Pubblicazione: (2025)
di: Yang, Chenghao, et al.
Pubblicazione: (2025)
Lossless Compression of Neural Network Components: Weights, Checkpoints, and K/V Caches in Low-Precision Formats
di: Heilper, Anat, et al.
Pubblicazione: (2025)
di: Heilper, Anat, et al.
Pubblicazione: (2025)
IterInject: Indirect Prompt Injection Against LLM Agents via Feedback-Guided Iterative Optimization
di: Chen, Zixuan, et al.
Pubblicazione: (2026)
di: Chen, Zixuan, et al.
Pubblicazione: (2026)
ProCut: LLM Prompt Compression via Attribution Estimation
di: Xu, Zhentao, et al.
Pubblicazione: (2025)
di: Xu, Zhentao, et al.
Pubblicazione: (2025)
All is Not Lost: LLM Recovery without Checkpoints
di: Blagoev, Nikolay, et al.
Pubblicazione: (2025)
di: Blagoev, Nikolay, et al.
Pubblicazione: (2025)
An Efficient Compression of Deep Neural Network Checkpoints Based on Prediction and Context Modeling
di: Kim, Yuriy, et al.
Pubblicazione: (2025)
di: Kim, Yuriy, et al.
Pubblicazione: (2025)
Context-Guided Spatial Feature Reconstruction for Efficient Semantic Segmentation
di: Ni, Zhenliang, et al.
Pubblicazione: (2024)
di: Ni, Zhenliang, et al.
Pubblicazione: (2024)
PolarQuant: Optimal Gaussian Weight Quantization via Hadamard Rotation for LLM Compression
di: Vicentino, Caio
Pubblicazione: (2026)
di: Vicentino, Caio
Pubblicazione: (2026)
LatentLLM: Attention-Aware Joint Tensor Compression
di: Koike-Akino, Toshiaki, et al.
Pubblicazione: (2025)
di: Koike-Akino, Toshiaki, et al.
Pubblicazione: (2025)
Compressed Decentralized Momentum Stochastic Gradient Methods for Nonconvex Optimization
di: Liu, Wei, et al.
Pubblicazione: (2025)
di: Liu, Wei, et al.
Pubblicazione: (2025)
Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
di: Dong, Peijie, et al.
Pubblicazione: (2025)
di: Dong, Peijie, et al.
Pubblicazione: (2025)
Predicting Human Mobility during Extreme Events via LLM-Enhanced Cross-City Learning
di: Tang, Yinzhou, et al.
Pubblicazione: (2025)
di: Tang, Yinzhou, et al.
Pubblicazione: (2025)
ExLLM: Experience-Enhanced LLM Optimization for Molecular Design and Beyond
di: Ran, Nian, et al.
Pubblicazione: (2025)
di: Ran, Nian, et al.
Pubblicazione: (2025)
Aggressive Compression Enables LLM Weight Theft
di: Brown, Davis, et al.
Pubblicazione: (2026)
di: Brown, Davis, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SlimLLM: Accurate Structured Pruning for Large Language Models
di: Guo, Jialong, et al.
Pubblicazione: (2025) -
PocketLLM: Ultimate Compression of Large Language Models via Meta Networks
di: Tian, Ye, et al.
Pubblicazione: (2025) -
LLM Data Selection and Utilization via Dynamic Bi-level Optimization
di: Yu, Yang, et al.
Pubblicazione: (2025) -
No Time to Waste: Squeeze Time into Channel for Mobile Video Understanding
di: Zhai, Yingjie, et al.
Pubblicazione: (2024) -
A Survey on Transformer Compression
di: Tang, Yehui, et al.
Pubblicazione: (2024)