Addition is almost all you need: Compressing large language models with double binary factorization
Fuente:
arXiv
Salvato in:
| Autori principali: | Boža, Vladimír, Macko, Vladimír |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Two Sparse Matrices are Better than One: Sparsifying Neural Networks with Double Sparse Factorization
di: Boža, Vladimír, et al.
Pubblicazione: (2024)
di: Boža, Vladimír, et al.
Pubblicazione: (2024)
MACKO: Sparse Matrix-Vector Multiplication for Low Sparsity
di: Macko, Vladimír, et al.
Pubblicazione: (2025)
di: Macko, Vladimír, et al.
Pubblicazione: (2025)
Attention and Compression is all you need for Controllably Efficient Language Models
di: Prakash, Jatin, et al.
Pubblicazione: (2025)
di: Prakash, Jatin, et al.
Pubblicazione: (2025)
Fast and Effective Weight Update for Pruned Large Language Models
di: Boža, Vladimír
Pubblicazione: (2024)
di: Boža, Vladimír
Pubblicazione: (2024)
One protein is all you need
di: Bushuiev, Anton, et al.
Pubblicazione: (2024)
di: Bushuiev, Anton, et al.
Pubblicazione: (2024)
Kolmogorov GAM Networks are all you need!
di: Polson, Sarah, et al.
Pubblicazione: (2025)
di: Polson, Sarah, et al.
Pubblicazione: (2025)
KV-weights are all you need for skipless transformers
di: Graef, Nils
Pubblicazione: (2024)
di: Graef, Nils
Pubblicazione: (2024)
Tabular Data: Is Deep Learning all you need?
di: Zabërgja, Guri, et al.
Pubblicazione: (2024)
di: Zabërgja, Guri, et al.
Pubblicazione: (2024)
Image compositing is all you need for data augmentation
di: Shermaine, Ang Jia Ning, et al.
Pubblicazione: (2025)
di: Shermaine, Ang Jia Ning, et al.
Pubblicazione: (2025)
Large Language Models aren't all that you need
di: Holla, Kiran Voderhobli, et al.
Pubblicazione: (2024)
di: Holla, Kiran Voderhobli, et al.
Pubblicazione: (2024)
Experts are all you need: A Composable Framework for Large Language Model Inference
di: Sridharan, Shrihari, et al.
Pubblicazione: (2025)
di: Sridharan, Shrihari, et al.
Pubblicazione: (2025)
Graph is all you need? Lightweight data-agnostic neural architecture search without training
di: Huang, Zhenhan, et al.
Pubblicazione: (2024)
di: Huang, Zhenhan, et al.
Pubblicazione: (2024)
Cross-Modal Safety Alignment: Is textual unlearning all you need?
di: Chakraborty, Trishna, et al.
Pubblicazione: (2024)
di: Chakraborty, Trishna, et al.
Pubblicazione: (2024)
Attention is all you need for boosting graph convolutional neural network
di: Wu, Yinwei
Pubblicazione: (2024)
di: Wu, Yinwei
Pubblicazione: (2024)
Linear attention is (maybe) all you need (to understand transformer optimization)
di: Ahn, Kwangjun, et al.
Pubblicazione: (2023)
di: Ahn, Kwangjun, et al.
Pubblicazione: (2023)
Simulation-based inference with scattering representations: scattering is all you need
di: Lin, Kiyam, et al.
Pubblicazione: (2024)
di: Lin, Kiyam, et al.
Pubblicazione: (2024)
DC is all you need: describing ReLU from a signal processing standpoint
di: Kechris, Christodoulos, et al.
Pubblicazione: (2024)
di: Kechris, Christodoulos, et al.
Pubblicazione: (2024)
Few Labels are all you need: A Weakly Supervised Framework for Appliance Localization in Smart-Meter Series
di: Petralia, Adrien, et al.
Pubblicazione: (2025)
di: Petralia, Adrien, et al.
Pubblicazione: (2025)
Neural Operator: Is data all you need to model the world? An insight into the paradigm of data-driven scientific ML
di: Viswanath, Hrishikesh, et al.
Pubblicazione: (2023)
di: Viswanath, Hrishikesh, et al.
Pubblicazione: (2023)
Attention is all you need for an improved CNN-based flash flood susceptibility modeling. The case of the ungauged Rheraya watershed, Morocco
di: Elghouat, Akram, et al.
Pubblicazione: (2024)
di: Elghouat, Akram, et al.
Pubblicazione: (2024)
Slim attention: cut your context memory in half without loss -- K-cache is all you need for MHA
di: Graef, Nils, et al.
Pubblicazione: (2025)
di: Graef, Nils, et al.
Pubblicazione: (2025)
Anti-concentration is (almost) all you need
di: Heinrich, Markus, et al.
Pubblicazione: (2025)
di: Heinrich, Markus, et al.
Pubblicazione: (2025)
Is attention all you need in medical image analysis? A review
di: Papanastasiou, Giorgos, et al.
Pubblicazione: (2023)
di: Papanastasiou, Giorgos, et al.
Pubblicazione: (2023)
Block removal for large language models through constrained binary optimization
di: Jansen, David, et al.
Pubblicazione: (2026)
di: Jansen, David, et al.
Pubblicazione: (2026)
1 bit is all we need: binary normalized neural networks
di: Cabral, Eduardo Lobo Lustoda, et al.
Pubblicazione: (2025)
di: Cabral, Eduardo Lobo Lustoda, et al.
Pubblicazione: (2025)
Why you don't overfit, and don't need Bayes if you only train for one epoch
di: Aitchison, Laurence
Pubblicazione: (2024)
di: Aitchison, Laurence
Pubblicazione: (2024)
Signformer is all you need: Towards Edge AI for Sign Language
di: Yang, Eta
Pubblicazione: (2024)
di: Yang, Eta
Pubblicazione: (2024)
Visual cognition in multimodal large language models
di: Buschoff, Luca M. Schulze, et al.
Pubblicazione: (2023)
di: Buschoff, Luca M. Schulze, et al.
Pubblicazione: (2023)
Hypothesis generation and updating in large language models
di: Xiong, Hua-Dong
Pubblicazione: (2026)
di: Xiong, Hua-Dong
Pubblicazione: (2026)
Representation in large language models
di: Yetman, Cameron
Pubblicazione: (2025)
di: Yetman, Cameron
Pubblicazione: (2025)
Compressed models are NOT miniature versions of large models
di: Rai, Rohit Raj, et al.
Pubblicazione: (2024)
di: Rai, Rohit Raj, et al.
Pubblicazione: (2024)
A Bayesian Optimization approach for calibrating large-scale activity-based transport models
di: Agriesti, Serio, et al.
Pubblicazione: (2023)
di: Agriesti, Serio, et al.
Pubblicazione: (2023)
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
di: Pacchiardi, Lorenzo, et al.
Pubblicazione: (2024)
di: Pacchiardi, Lorenzo, et al.
Pubblicazione: (2024)
Amortizing intractable inference in large language models
di: Hu, Edward J., et al.
Pubblicazione: (2023)
di: Hu, Edward J., et al.
Pubblicazione: (2023)
Layer-wise dynamic rank for compressing large language models
di: Mi, Zhendong, et al.
Pubblicazione: (2025)
di: Mi, Zhendong, et al.
Pubblicazione: (2025)
NIRVANA: Structured pruning reimagined for large language models compression
di: Ai, Mengting, et al.
Pubblicazione: (2025)
di: Ai, Mengting, et al.
Pubblicazione: (2025)
Training microrobots to swim by a large language model
di: Xu, Zhuoqun, et al.
Pubblicazione: (2024)
di: Xu, Zhuoqun, et al.
Pubblicazione: (2024)
Alignment faking in large language models
di: Greenblatt, Ryan, et al.
Pubblicazione: (2024)
di: Greenblatt, Ryan, et al.
Pubblicazione: (2024)
AI-AI Bias: large language models favor communications generated by large language models
di: Laurito, Walter, et al.
Pubblicazione: (2024)
di: Laurito, Walter, et al.
Pubblicazione: (2024)
Residual vector quantization for KV cache compression in large language model
di: Kumar, Ankur
Pubblicazione: (2024)
di: Kumar, Ankur
Pubblicazione: (2024)
Documenti analoghi
-
Two Sparse Matrices are Better than One: Sparsifying Neural Networks with Double Sparse Factorization
di: Boža, Vladimír, et al.
Pubblicazione: (2024) -
MACKO: Sparse Matrix-Vector Multiplication for Low Sparsity
di: Macko, Vladimír, et al.
Pubblicazione: (2025) -
Attention and Compression is all you need for Controllably Efficient Language Models
di: Prakash, Jatin, et al.
Pubblicazione: (2025) -
Fast and Effective Weight Update for Pruned Large Language Models
di: Boža, Vladimír
Pubblicazione: (2024) -
One protein is all you need
di: Bushuiev, Anton, et al.
Pubblicazione: (2024)