Large Language Model Compression via the Nested Activation-Aware Decomposition
Fuente:
arXiv
Salvato in:
| Autori principali: | Lu, Jun, Xu, Tianyi, Ding, Bill, Li, David, Kang, Yu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Improving embedding with contrastive fine-tuning on small datasets with expert-augmented scores
di: Lu, Jun, et al.
Pubblicazione: (2024)
di: Lu, Jun, et al.
Pubblicazione: (2024)
CURing Large Models: Compression via CUR Decomposition
di: Park, Sanghyeon, et al.
Pubblicazione: (2025)
di: Park, Sanghyeon, et al.
Pubblicazione: (2025)
Entropy Regularizing Activation: Boosting Continuous Control, Large Language Models, and Image Classification with Activation as Entropy Constraints
di: Kang, Zilin, et al.
Pubblicazione: (2025)
di: Kang, Zilin, et al.
Pubblicazione: (2025)
SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression
di: Wang, Xin, et al.
Pubblicazione: (2024)
di: Wang, Xin, et al.
Pubblicazione: (2024)
Activation Sparsity Opportunities for Compressing General Large Language Models
di: Dhar, Nobel, et al.
Pubblicazione: (2024)
di: Dhar, Nobel, et al.
Pubblicazione: (2024)
BALF: Budgeted Activation-Aware Low-Rank Factorization for Fine-Tuning-Free Model Compression
di: González-Martínez, David
Pubblicazione: (2025)
di: González-Martínez, David
Pubblicazione: (2025)
SLaB: Sparse-Lowrank-Binary Decomposition for Efficient Large Language Models
di: Li, Ziwei, et al.
Pubblicazione: (2026)
di: Li, Ziwei, et al.
Pubblicazione: (2026)
Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language Models
di: Leask, Patrick, et al.
Pubblicazione: (2025)
di: Leask, Patrick, et al.
Pubblicazione: (2025)
TensorGPT: Efficient Compression of Large Language Models based on Tensor-Train Decomposition
di: Xu, Mingxue, et al.
Pubblicazione: (2023)
di: Xu, Mingxue, et al.
Pubblicazione: (2023)
Sparse Gradient Compression for Fine-Tuning Large Language Models
di: Yang, David H., et al.
Pubblicazione: (2025)
di: Yang, David H., et al.
Pubblicazione: (2025)
Compressing Large Language Models using Low Rank and Low Precision Decomposition
di: Saha, Rajarshi, et al.
Pubblicazione: (2024)
di: Saha, Rajarshi, et al.
Pubblicazione: (2024)
Deep Hierarchical Learning with Nested Subspace Networks for Large Language Models
di: Rauba, Paulius, et al.
Pubblicazione: (2025)
di: Rauba, Paulius, et al.
Pubblicazione: (2025)
ELAS: Efficient Pre-Training of Low-Rank Large Language Models via 2:4 Activation Sparsity
di: Li, Jiaxi, et al.
Pubblicazione: (2026)
di: Li, Jiaxi, et al.
Pubblicazione: (2026)
AgentCompress: Task-Aware Compression for Affordable Large Language Model Agents
di: Taha, Zuhair Ahmed Khan, et al.
Pubblicazione: (2026)
di: Taha, Zuhair Ahmed Khan, et al.
Pubblicazione: (2026)
CALR: Corrective Adaptive Low-Rank Decomposition for Efficient Large Language Model Layer Compression
di: Kautsar, Muchammad Daniyal, et al.
Pubblicazione: (2025)
di: Kautsar, Muchammad Daniyal, et al.
Pubblicazione: (2025)
FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model Deployment
di: Zaccone, Riccardo, et al.
Pubblicazione: (2026)
di: Zaccone, Riccardo, et al.
Pubblicazione: (2026)
Layer as Puzzle Pieces: Compressing Large Language Models through Layer Concatenation
di: Wang, Fei, et al.
Pubblicazione: (2025)
di: Wang, Fei, et al.
Pubblicazione: (2025)
WINA: Weight Informed Neuron Activation for Accelerating Large Language Model Inference
di: Chen, Sihan, et al.
Pubblicazione: (2025)
di: Chen, Sihan, et al.
Pubblicazione: (2025)
Activation Map Compression through Tensor Decomposition for Deep Learning
di: Nguyen, Le-Trung, et al.
Pubblicazione: (2024)
di: Nguyen, Le-Trung, et al.
Pubblicazione: (2024)
HierRouter: Coordinated Routing of Specialized Large Language Models via Reinforcement Learning
di: Gupta, Nikunj, et al.
Pubblicazione: (2025)
di: Gupta, Nikunj, et al.
Pubblicazione: (2025)
Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
di: Mao, Yu, et al.
Pubblicazione: (2025)
di: Mao, Yu, et al.
Pubblicazione: (2025)
To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration
di: Yang, Zeyu, et al.
Pubblicazione: (2025)
di: Yang, Zeyu, et al.
Pubblicazione: (2025)
Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression
di: Tang, Yuntian, et al.
Pubblicazione: (2026)
di: Tang, Yuntian, et al.
Pubblicazione: (2026)
Matrix Decomposition and Applications
di: Lu, Jun
Pubblicazione: (2022)
di: Lu, Jun
Pubblicazione: (2022)
On the Compressibility of Quantized Large Language Models
di: Mao, Yu, et al.
Pubblicazione: (2024)
di: Mao, Yu, et al.
Pubblicazione: (2024)
Extreme Compression of Large Language Models via Additive Quantization
di: Egiazarian, Vage, et al.
Pubblicazione: (2024)
di: Egiazarian, Vage, et al.
Pubblicazione: (2024)
MoDeGPT: Modular Decomposition for Large Language Model Compression
di: Lin, Chi-Heng, et al.
Pubblicazione: (2024)
di: Lin, Chi-Heng, et al.
Pubblicazione: (2024)
Capability-Guided Compression: Toward Interpretability-Aware Budget Allocation for Large Language Models
di: Gupta, Rishaank
Pubblicazione: (2026)
di: Gupta, Rishaank
Pubblicazione: (2026)
NestQuant: Post-Training Integer-Nesting Quantization for On-Device DNN
di: Xie, Jianhang, et al.
Pubblicazione: (2025)
di: Xie, Jianhang, et al.
Pubblicazione: (2025)
Importance-Guided Basis Selection for Low-Rank Decomposition of Large Language Models
di: Asante, Daniel Agyei, et al.
Pubblicazione: (2026)
di: Asante, Daniel Agyei, et al.
Pubblicazione: (2026)
KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
FastMMoE: Accelerating Multimodal Large Language Models through Dynamic Expert Activation and Routing-Aware Token Pruning
di: Xia, Guoyang, et al.
Pubblicazione: (2025)
di: Xia, Guoyang, et al.
Pubblicazione: (2025)
Large Language Models for Anomaly and Out-of-Distribution Detection: A Survey
di: Xu, Ruiyao, et al.
Pubblicazione: (2024)
di: Xu, Ruiyao, et al.
Pubblicazione: (2024)
Evo: Autoregressive-Diffusion Large Language Models with Evolving Balance
di: Wu, Junde, et al.
Pubblicazione: (2026)
di: Wu, Junde, et al.
Pubblicazione: (2026)
Compression Aware Certified Training
di: Xu, Changming, et al.
Pubblicazione: (2025)
di: Xu, Changming, et al.
Pubblicazione: (2025)
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
di: Zhao, Youpeng, et al.
Pubblicazione: (2024)
di: Zhao, Youpeng, et al.
Pubblicazione: (2024)
LeanQuant: Accurate and Scalable Large Language Model Quantization with Loss-error-aware Grid
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
Basis Sharing: Cross-Layer Parameter Sharing for Large Language Model Compression
di: Wang, Jingcun, et al.
Pubblicazione: (2024)
di: Wang, Jingcun, et al.
Pubblicazione: (2024)
Shuttle Between the Instructions and the Parameters of Large Language Models
di: Sun, Wangtao, et al.
Pubblicazione: (2025)
di: Sun, Wangtao, et al.
Pubblicazione: (2025)
ESPACE: Dimensionality Reduction of Activations for Model Compression
di: Sakr, Charbel, et al.
Pubblicazione: (2024)
di: Sakr, Charbel, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Improving embedding with contrastive fine-tuning on small datasets with expert-augmented scores
di: Lu, Jun, et al.
Pubblicazione: (2024) -
CURing Large Models: Compression via CUR Decomposition
di: Park, Sanghyeon, et al.
Pubblicazione: (2025) -
Entropy Regularizing Activation: Boosting Continuous Control, Large Language Models, and Image Classification with Activation as Entropy Constraints
di: Kang, Zilin, et al.
Pubblicazione: (2025) -
SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression
di: Wang, Xin, et al.
Pubblicazione: (2024) -
Activation Sparsity Opportunities for Compressing General Large Language Models
di: Dhar, Nobel, et al.
Pubblicazione: (2024)