Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Ilin, Ivan, Richtarik, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Teaching and Learning under Deductive Errors
by: Telle, Jan Arne, et al.
Published: (2026)
by: Telle, Jan Arne, et al.
Published: (2026)
Autoencoded UMAP-Enhanced Clustering for Unsupervised Learning
by: Chavooshi, Malihehsadat, et al.
Published: (2025)
by: Chavooshi, Malihehsadat, et al.
Published: (2025)
Binarized Neural Networks Converge Toward Algorithmic Simplicity: Empirical Support for the Learning-as-Compression Hypothesis
by: Sakabe, Eduardo Y., et al.
Published: (2025)
by: Sakabe, Eduardo Y., et al.
Published: (2025)
Neural Network Approximation: A View from Polytope Decomposition
by: Li, ZeYu, et al.
Published: (2026)
by: Li, ZeYu, et al.
Published: (2026)
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization
by: Ustaomeroglu, Muhammed, et al.
Published: (2025)
by: Ustaomeroglu, Muhammed, et al.
Published: (2025)
Backpropagation Through Time For Networks With Long-Term Dependencies
by: Bird, George, et al.
Published: (2021)
by: Bird, George, et al.
Published: (2021)
Hessian of Perplexity for Large Language Models by PyTorch autograd (Open Source)
by: Ilin, Ivan
Published: (2025)
by: Ilin, Ivan
Published: (2025)
Grouped Sequential Optimization Strategy -- the Application of Hyperparameter Importance Assessment in Deep Learning
by: Wang, Ruinan, et al.
Published: (2025)
by: Wang, Ruinan, et al.
Published: (2025)
Recent Advances in Named Entity Recognition: A Comprehensive Survey and Comparative Study
by: Keraghel, Imed, et al.
Published: (2024)
by: Keraghel, Imed, et al.
Published: (2024)
InhibiDistilbert: Knowledge Distillation for a ReLU and Addition-based Transformer
by: Zhang, Tony, et al.
Published: (2025)
by: Zhang, Tony, et al.
Published: (2025)
Modularity in Transformers: Investigating Neuron Separability & Specialization
by: Pochinkov, Nicholas, et al.
Published: (2024)
by: Pochinkov, Nicholas, et al.
Published: (2024)
The two clocks and the innovation window: When and how generative models learn rules
by: Wang, Binxu, et al.
Published: (2026)
by: Wang, Binxu, et al.
Published: (2026)
On the Origin of Algorithmic Progress in AI
by: Gundlach, Hans, et al.
Published: (2025)
by: Gundlach, Hans, et al.
Published: (2025)
A Language Model-Driven Semi-Supervised Ensemble Framework for Illicit Market Detection Across Deep/Dark Web and Social Platforms
by: Yazdanjue, Navid, et al.
Published: (2025)
by: Yazdanjue, Navid, et al.
Published: (2025)
Learning Neural Network Classifiers with Low Model Complexity
by: Jayadeva, et al.
Published: (2017)
by: Jayadeva, et al.
Published: (2017)
Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization
by: Li, Yin
Published: (2025)
by: Li, Yin
Published: (2025)
A Special Case of Quadratic Extrapolation Under the Neural Tangent Kernel
by: Kim, Abiel
Published: (2025)
by: Kim, Abiel
Published: (2025)
Greedy feature selection: Classifier-dependent feature selection via greedy methods
by: Camattari, Fabiana, et al.
Published: (2024)
by: Camattari, Fabiana, et al.
Published: (2024)
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
by: Fu, Tianyu, et al.
Published: (2025)
by: Fu, Tianyu, et al.
Published: (2025)
Understanding the Nature of Generative AI as Threshold Logic in High-Dimensional Space
by: Levin, Ilya
Published: (2026)
by: Levin, Ilya
Published: (2026)
OptPO: Optimal Rollout Allocation for Test-time Policy Optimization
by: Wang, Youkang, et al.
Published: (2025)
by: Wang, Youkang, et al.
Published: (2025)
ADAPT: Lightweight, Long-Range Machine Learning Force Fields Without Graphs
by: Dramko, Evan, et al.
Published: (2025)
by: Dramko, Evan, et al.
Published: (2025)
FSD-CAP: Fractional Subgraph Diffusion with Class-Aware Propagation for Graph Feature Imputation
by: Qiao, Xin, et al.
Published: (2026)
by: Qiao, Xin, et al.
Published: (2026)
Parameter-Efficient Transformer Embeddings
by: Ndubuaku, Henry, et al.
Published: (2025)
by: Ndubuaku, Henry, et al.
Published: (2025)
QGraphLIME - Explaining Quantum Graph Neural Networks
by: Jena, Haribandhu, et al.
Published: (2025)
by: Jena, Haribandhu, et al.
Published: (2025)
Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization
by: Pavlov, Gorgi
Published: (2026)
by: Pavlov, Gorgi
Published: (2026)
High-arity Sample Compression
by: Coregliano, Leonardo N., et al.
Published: (2026)
by: Coregliano, Leonardo N., et al.
Published: (2026)
Different Statistical Perspectives for Understanding Generalisation in Graph Neural Networks
by: Ayday, Nil, et al.
Published: (2026)
by: Ayday, Nil, et al.
Published: (2026)
Stick to your Role! Stability of Personal Values Expressed in Large Language Models
by: Kovač, Grgur, et al.
Published: (2024)
by: Kovač, Grgur, et al.
Published: (2024)
Golden Handcuffs make safer AI agents
by: Ebtekar, Aram, et al.
Published: (2026)
by: Ebtekar, Aram, et al.
Published: (2026)
On Halting vs Converging in Recurrent Graph Neural Networks
by: Bollen, Jeroen, et al.
Published: (2026)
by: Bollen, Jeroen, et al.
Published: (2026)
Bounds on the Generalization Error in Active Learning
by: Menden, Vincent, et al.
Published: (2024)
by: Menden, Vincent, et al.
Published: (2024)
Swap Agnostic Learning, or Characterizing Omniprediction via Multicalibration
by: Gopalan, Parikshit, et al.
Published: (2023)
by: Gopalan, Parikshit, et al.
Published: (2023)
The Price of Robustness: Stable Classifiers Need Overparameterization
by: von Berg, Jonas, et al.
Published: (2026)
by: von Berg, Jonas, et al.
Published: (2026)
Enhanced QKNorm normalization for neural transformers with the Lp norm
by: Lopez-Rubio, Ezequiel, et al.
Published: (2026)
by: Lopez-Rubio, Ezequiel, et al.
Published: (2026)
A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness
by: Wang, Fali, et al.
Published: (2024)
by: Wang, Fali, et al.
Published: (2024)
Smoothed Embeddings for Robust Language Models
by: Hase, Ryo, et al.
Published: (2025)
by: Hase, Ryo, et al.
Published: (2025)
DYNAMAX: Dynamic computing for Transformers and Mamba based architectures
by: Nogales, Miguel, et al.
Published: (2025)
by: Nogales, Miguel, et al.
Published: (2025)
Error Bounds for Learning with Vector-Valued Random Features
by: Lanthaler, Samuel, et al.
Published: (2023)
by: Lanthaler, Samuel, et al.
Published: (2023)
Beyond Discreteness: Sample Complexity Analysis of Straight-Through Estimator for 1-bit Quantization
by: Jeong, Halyun, et al.
Published: (2025)
by: Jeong, Halyun, et al.
Published: (2025)
Similar Items
-
Teaching and Learning under Deductive Errors
by: Telle, Jan Arne, et al.
Published: (2026) -
Autoencoded UMAP-Enhanced Clustering for Unsupervised Learning
by: Chavooshi, Malihehsadat, et al.
Published: (2025) -
Binarized Neural Networks Converge Toward Algorithmic Simplicity: Empirical Support for the Learning-as-Compression Hypothesis
by: Sakabe, Eduardo Y., et al.
Published: (2025) -
Neural Network Approximation: A View from Polytope Decomposition
by: Li, ZeYu, et al.
Published: (2026) -
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization
by: Ustaomeroglu, Muhammed, et al.
Published: (2025)