Addition is almost all you need: Compressing large language models with double binary factorization
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Boža, Vladimír, Macko, Vladimír |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Two Sparse Matrices are Better than One: Sparsifying Neural Networks with Double Sparse Factorization
par: Boža, Vladimír, et autres
Publié: (2024)
par: Boža, Vladimír, et autres
Publié: (2024)
MACKO: Sparse Matrix-Vector Multiplication for Low Sparsity
par: Macko, Vladimír, et autres
Publié: (2025)
par: Macko, Vladimír, et autres
Publié: (2025)
Attention and Compression is all you need for Controllably Efficient Language Models
par: Prakash, Jatin, et autres
Publié: (2025)
par: Prakash, Jatin, et autres
Publié: (2025)
Fast and Effective Weight Update for Pruned Large Language Models
par: Boža, Vladimír
Publié: (2024)
par: Boža, Vladimír
Publié: (2024)
One protein is all you need
par: Bushuiev, Anton, et autres
Publié: (2024)
par: Bushuiev, Anton, et autres
Publié: (2024)
Kolmogorov GAM Networks are all you need!
par: Polson, Sarah, et autres
Publié: (2025)
par: Polson, Sarah, et autres
Publié: (2025)
KV-weights are all you need for skipless transformers
par: Graef, Nils
Publié: (2024)
par: Graef, Nils
Publié: (2024)
Tabular Data: Is Deep Learning all you need?
par: Zabërgja, Guri, et autres
Publié: (2024)
par: Zabërgja, Guri, et autres
Publié: (2024)
Image compositing is all you need for data augmentation
par: Shermaine, Ang Jia Ning, et autres
Publié: (2025)
par: Shermaine, Ang Jia Ning, et autres
Publié: (2025)
Large Language Models aren't all that you need
par: Holla, Kiran Voderhobli, et autres
Publié: (2024)
par: Holla, Kiran Voderhobli, et autres
Publié: (2024)
Experts are all you need: A Composable Framework for Large Language Model Inference
par: Sridharan, Shrihari, et autres
Publié: (2025)
par: Sridharan, Shrihari, et autres
Publié: (2025)
Graph is all you need? Lightweight data-agnostic neural architecture search without training
par: Huang, Zhenhan, et autres
Publié: (2024)
par: Huang, Zhenhan, et autres
Publié: (2024)
Cross-Modal Safety Alignment: Is textual unlearning all you need?
par: Chakraborty, Trishna, et autres
Publié: (2024)
par: Chakraborty, Trishna, et autres
Publié: (2024)
Attention is all you need for boosting graph convolutional neural network
par: Wu, Yinwei
Publié: (2024)
par: Wu, Yinwei
Publié: (2024)
Linear attention is (maybe) all you need (to understand transformer optimization)
par: Ahn, Kwangjun, et autres
Publié: (2023)
par: Ahn, Kwangjun, et autres
Publié: (2023)
Simulation-based inference with scattering representations: scattering is all you need
par: Lin, Kiyam, et autres
Publié: (2024)
par: Lin, Kiyam, et autres
Publié: (2024)
DC is all you need: describing ReLU from a signal processing standpoint
par: Kechris, Christodoulos, et autres
Publié: (2024)
par: Kechris, Christodoulos, et autres
Publié: (2024)
Few Labels are all you need: A Weakly Supervised Framework for Appliance Localization in Smart-Meter Series
par: Petralia, Adrien, et autres
Publié: (2025)
par: Petralia, Adrien, et autres
Publié: (2025)
Neural Operator: Is data all you need to model the world? An insight into the paradigm of data-driven scientific ML
par: Viswanath, Hrishikesh, et autres
Publié: (2023)
par: Viswanath, Hrishikesh, et autres
Publié: (2023)
Attention is all you need for an improved CNN-based flash flood susceptibility modeling. The case of the ungauged Rheraya watershed, Morocco
par: Elghouat, Akram, et autres
Publié: (2024)
par: Elghouat, Akram, et autres
Publié: (2024)
Slim attention: cut your context memory in half without loss -- K-cache is all you need for MHA
par: Graef, Nils, et autres
Publié: (2025)
par: Graef, Nils, et autres
Publié: (2025)
Anti-concentration is (almost) all you need
par: Heinrich, Markus, et autres
Publié: (2025)
par: Heinrich, Markus, et autres
Publié: (2025)
Is attention all you need in medical image analysis? A review
par: Papanastasiou, Giorgos, et autres
Publié: (2023)
par: Papanastasiou, Giorgos, et autres
Publié: (2023)
Block removal for large language models through constrained binary optimization
par: Jansen, David, et autres
Publié: (2026)
par: Jansen, David, et autres
Publié: (2026)
1 bit is all we need: binary normalized neural networks
par: Cabral, Eduardo Lobo Lustoda, et autres
Publié: (2025)
par: Cabral, Eduardo Lobo Lustoda, et autres
Publié: (2025)
Why you don't overfit, and don't need Bayes if you only train for one epoch
par: Aitchison, Laurence
Publié: (2024)
par: Aitchison, Laurence
Publié: (2024)
Signformer is all you need: Towards Edge AI for Sign Language
par: Yang, Eta
Publié: (2024)
par: Yang, Eta
Publié: (2024)
Visual cognition in multimodal large language models
par: Buschoff, Luca M. Schulze, et autres
Publié: (2023)
par: Buschoff, Luca M. Schulze, et autres
Publié: (2023)
Hypothesis generation and updating in large language models
par: Xiong, Hua-Dong
Publié: (2026)
par: Xiong, Hua-Dong
Publié: (2026)
Representation in large language models
par: Yetman, Cameron
Publié: (2025)
par: Yetman, Cameron
Publié: (2025)
Compressed models are NOT miniature versions of large models
par: Rai, Rohit Raj, et autres
Publié: (2024)
par: Rai, Rohit Raj, et autres
Publié: (2024)
A Bayesian Optimization approach for calibrating large-scale activity-based transport models
par: Agriesti, Serio, et autres
Publié: (2023)
par: Agriesti, Serio, et autres
Publié: (2023)
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
par: Pacchiardi, Lorenzo, et autres
Publié: (2024)
par: Pacchiardi, Lorenzo, et autres
Publié: (2024)
Amortizing intractable inference in large language models
par: Hu, Edward J., et autres
Publié: (2023)
par: Hu, Edward J., et autres
Publié: (2023)
Layer-wise dynamic rank for compressing large language models
par: Mi, Zhendong, et autres
Publié: (2025)
par: Mi, Zhendong, et autres
Publié: (2025)
NIRVANA: Structured pruning reimagined for large language models compression
par: Ai, Mengting, et autres
Publié: (2025)
par: Ai, Mengting, et autres
Publié: (2025)
Training microrobots to swim by a large language model
par: Xu, Zhuoqun, et autres
Publié: (2024)
par: Xu, Zhuoqun, et autres
Publié: (2024)
Alignment faking in large language models
par: Greenblatt, Ryan, et autres
Publié: (2024)
par: Greenblatt, Ryan, et autres
Publié: (2024)
AI-AI Bias: large language models favor communications generated by large language models
par: Laurito, Walter, et autres
Publié: (2024)
par: Laurito, Walter, et autres
Publié: (2024)
Residual vector quantization for KV cache compression in large language model
par: Kumar, Ankur
Publié: (2024)
par: Kumar, Ankur
Publié: (2024)
Documents similaires
-
Two Sparse Matrices are Better than One: Sparsifying Neural Networks with Double Sparse Factorization
par: Boža, Vladimír, et autres
Publié: (2024) -
MACKO: Sparse Matrix-Vector Multiplication for Low Sparsity
par: Macko, Vladimír, et autres
Publié: (2025) -
Attention and Compression is all you need for Controllably Efficient Language Models
par: Prakash, Jatin, et autres
Publié: (2025) -
Fast and Effective Weight Update for Pruned Large Language Models
par: Boža, Vladimír
Publié: (2024) -
One protein is all you need
par: Bushuiev, Anton, et autres
Publié: (2024)