To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Zeyu, Zhang, Tianyi, Xie, Jianwen, Li, Chuan, Xu, Zhaozhuo, Shrivastava, Anshumali |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Scout Before You Attend: Sketch-and-Walk Sparse Attention for Efficient LLM Inference
di: Le, Hoang Anh Duy, et al.
Pubblicazione: (2026)
di: Le, Hoang Anh Duy, et al.
Pubblicazione: (2026)
Superintelligent Retrieval Agent: The Next Frontier of Information Retrieval
di: Yang, Zeyu, et al.
Pubblicazione: (2026)
di: Yang, Zeyu, et al.
Pubblicazione: (2026)
NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
FineZip : Pushing the Limits of Large Language Models for Practical Lossless Text Compression
di: Mittu, Fazal, et al.
Pubblicazione: (2024)
di: Mittu, Fazal, et al.
Pubblicazione: (2024)
Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era
di: Li, Dawei, et al.
Pubblicazione: (2025)
di: Li, Dawei, et al.
Pubblicazione: (2025)
DP-CSGP: Differentially Private Stochastic Gradient Push with Compressed Communication
di: Zhu, Zehan, et al.
Pubblicazione: (2025)
di: Zhu, Zehan, et al.
Pubblicazione: (2025)
LeanQuant: Accurate and Scalable Large Language Model Quantization with Loss-error-aware Grid
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
OAT-Rephrase: Optimization-Aware Training Data Rephrasing for Zeroth-Order LLM Fine-Tuning
di: Long, Jikai, et al.
Pubblicazione: (2025)
di: Long, Jikai, et al.
Pubblicazione: (2025)
70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)
di: Zhang, Tianyi, et al.
Pubblicazione: (2025)
di: Zhang, Tianyi, et al.
Pubblicazione: (2025)
Lossless Compression: A New Benchmark for Time Series Model Evaluation
di: Wan, Meng, et al.
Pubblicazione: (2025)
di: Wan, Meng, et al.
Pubblicazione: (2025)
Lossless Model Compression via Joint Low-Rank Factorization Optimization
di: Zhang, Boyang, et al.
Pubblicazione: (2024)
di: Zhang, Boyang, et al.
Pubblicazione: (2024)
Randomized Antipodal Search Done Right for Data Pareto Improvement of LLM Unlearning
di: Liu, Ziwen, et al.
Pubblicazione: (2026)
di: Liu, Ziwen, et al.
Pubblicazione: (2026)
Lossless Token Sequence Compression via Meta-Tokens
di: Harvill, John, et al.
Pubblicazione: (2025)
di: Harvill, John, et al.
Pubblicazione: (2025)
Approaches to Responsible Governance of GenAI in Organizations
di: Gandhi, Dhari, et al.
Pubblicazione: (2025)
di: Gandhi, Dhari, et al.
Pubblicazione: (2025)
Lossless KV Cache Compression to 2%
di: Yang, Zhen, et al.
Pubblicazione: (2024)
di: Yang, Zhen, et al.
Pubblicazione: (2024)
Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
di: Long, Phillip, et al.
Pubblicazione: (2026)
di: Long, Phillip, et al.
Pubblicazione: (2026)
Sketch to Adapt: Fine-Tunable Sketches for Efficient LLM Adaptation
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
Lossless Compression of Neural Network Components: Weights, Checkpoints, and K/V Caches in Low-Precision Formats
di: Heilper, Anat, et al.
Pubblicazione: (2025)
di: Heilper, Anat, et al.
Pubblicazione: (2025)
AlphaZip: Neural Network-Enhanced Lossless Text Compression
di: Narashiman, Swathi Shree, et al.
Pubblicazione: (2024)
di: Narashiman, Swathi Shree, et al.
Pubblicazione: (2024)
TurboAngle: Near-Lossless KV Cache Compression via Uniform Angle Quantization
di: Patel, Dipkumar
Pubblicazione: (2026)
di: Patel, Dipkumar
Pubblicazione: (2026)
DEL-ToM: Inference-Time Scaling for Theory-of-Mind Reasoning via Dynamic Epistemic Logic
di: Wu, Yuheng, et al.
Pubblicazione: (2025)
di: Wu, Yuheng, et al.
Pubblicazione: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
di: Tian, Yuxuan, et al.
Pubblicazione: (2025)
di: Tian, Yuxuan, et al.
Pubblicazione: (2025)
RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
di: Joshi, Sahil, et al.
Pubblicazione: (2025)
di: Joshi, Sahil, et al.
Pubblicazione: (2025)
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
di: Kang, Hao, et al.
Pubblicazione: (2024)
di: Kang, Hao, et al.
Pubblicazione: (2024)
Near-Lossless Model Compression Enables Longer Context Inference in DNA Large Language Models
di: Zhu, Rui, et al.
Pubblicazione: (2025)
di: Zhu, Rui, et al.
Pubblicazione: (2025)
CoVE: Compressed Vocabulary Expansion Makes Better LLM-based Recommender Systems
di: Zhang, Haochen, et al.
Pubblicazione: (2025)
di: Zhang, Haochen, et al.
Pubblicazione: (2025)
GenAI-FDIA: Physics-Informed Generative Models for False Data Injection Attacks
di: Razzaque, Mohammad A., et al.
Pubblicazione: (2026)
di: Razzaque, Mohammad A., et al.
Pubblicazione: (2026)
DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
di: Yu, Xiaoming, et al.
Pubblicazione: (2026)
di: Yu, Xiaoming, et al.
Pubblicazione: (2026)
In the Mood to Exclude: Revitalizing Trespass to Chattels in the Era of GenAI Scraping
di: Atkinson, David
Pubblicazione: (2025)
di: Atkinson, David
Pubblicazione: (2025)
Compression for Better: A General and Stable Lossless Compression Framework
di: Zhang, Boyang, et al.
Pubblicazione: (2024)
di: Zhang, Boyang, et al.
Pubblicazione: (2024)
Open Foundation Models in Healthcare: Challenges, Paradoxes, and Opportunities with GenAI Driven Personalized Prescription
di: Alkaeed, Mahdi, et al.
Pubblicazione: (2025)
di: Alkaeed, Mahdi, et al.
Pubblicazione: (2025)
FTA generation using GenAI with an Autonomy sensor Usecase
di: Shetiya, Sneha Sudhir, et al.
Pubblicazione: (2024)
di: Shetiya, Sneha Sudhir, et al.
Pubblicazione: (2024)
Test-Time Steering for Lossless Text Compression via Weighted Product of Experts
di: Zhang, Qihang, et al.
Pubblicazione: (2025)
di: Zhang, Qihang, et al.
Pubblicazione: (2025)
REFRAG: Rethinking RAG based Decoding
di: Lin, Xiaoqiang, et al.
Pubblicazione: (2025)
di: Lin, Xiaoqiang, et al.
Pubblicazione: (2025)
How to Strategize Human Content Creation in the Era of GenAI?
di: Esmaeili, Seyed A., et al.
Pubblicazione: (2024)
di: Esmaeili, Seyed A., et al.
Pubblicazione: (2024)
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
di: Guo, Yipin, et al.
Pubblicazione: (2026)
di: Guo, Yipin, et al.
Pubblicazione: (2026)
Hyper-Compression: Model Compression via Hyperfunction
di: Fan, Fenglei, et al.
Pubblicazione: (2024)
di: Fan, Fenglei, et al.
Pubblicazione: (2024)
When Reasoning Meets Compression: Understanding the Effects of LLMs Compression on Large Reasoning Models
di: Zhang, Nan, et al.
Pubblicazione: (2025)
di: Zhang, Nan, et al.
Pubblicazione: (2025)
Industrializing Prediction-Powered Inference: The GLIDE Library for Reliable GenAI and Agentic Systems Evaluation
di: Martinon, Grégoire, et al.
Pubblicazione: (2026)
di: Martinon, Grégoire, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Scout Before You Attend: Sketch-and-Walk Sparse Attention for Efficient LLM Inference
di: Le, Hoang Anh Duy, et al.
Pubblicazione: (2026) -
Superintelligent Retrieval Agent: The Next Frontier of Information Retrieval
di: Yang, Zeyu, et al.
Pubblicazione: (2026) -
NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention
di: Zhang, Tianyi, et al.
Pubblicazione: (2024) -
KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization
di: Zhang, Tianyi, et al.
Pubblicazione: (2024) -
FineZip : Pushing the Limits of Large Language Models for Practical Lossless Text Compression
di: Mittu, Fazal, et al.
Pubblicazione: (2024)