Quantization of Large Language Models with an Overdetermined Basis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Merkulov, Daniil, Cherniuk, Daria, Rudikov, Alexander, Oseledets, Ivan, Muravleva, Ekaterina, Mikhalev, Aleksandr, Kashin, Boris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Run LoRA Run: Faster and Lighter LoRA Implementations
von: Cherniuk, Daria, et al.
Veröffentlicht: (2023)
von: Cherniuk, Daria, et al.
Veröffentlicht: (2023)
LoTR: Low Tensor Rank Weight Adaptation
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2024)
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2024)
Deep Learning for Subspace Regression
von: Fanaskov, Vladimir, et al.
Veröffentlicht: (2025)
von: Fanaskov, Vladimir, et al.
Veröffentlicht: (2025)
Bayesian Inverse Problems Meet Flow Matching: Efficient and Flexible Inference via Transformers
von: Sherki, Daniil, et al.
Veröffentlicht: (2025)
von: Sherki, Daniil, et al.
Veröffentlicht: (2025)
ConDiff: A Challenging Dataset for Neural Solvers of Partial Differential Equations
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2024)
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2024)
Learning from Linear Algebra: A Graph Neural Network Approach to Preconditioner Design for Conjugate Gradient Solvers
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2024)
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2024)
Spectral Analysis of the Weighted Frobenius Objective
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2025)
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2025)
Message-Passing GNNs Fail to Approximate Sparse Triangular Factorizations
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2025)
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2025)
NNTile: a machine learning framework capable of training extremely large GPT language models on a single node
von: Mikhalev, Aleksandr, et al.
Veröffentlicht: (2025)
von: Mikhalev, Aleksandr, et al.
Veröffentlicht: (2025)
Astral: training physics-informed neural networks with error majorants
von: Fanaskov, Vladimir, et al.
Veröffentlicht: (2024)
von: Fanaskov, Vladimir, et al.
Veröffentlicht: (2024)
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
von: Li, Pengyi, et al.
Veröffentlicht: (2025)
von: Li, Pengyi, et al.
Veröffentlicht: (2025)
Generalized Fisher-Weighted SVD: Scalable Kronecker-Factored Fisher Approximation for Compressing Large Language Models
von: Chekalina, Viktoriia, et al.
Veröffentlicht: (2025)
von: Chekalina, Viktoriia, et al.
Veröffentlicht: (2025)
OpenAutoNLU: Open Source AutoML Library for NLU
von: Arshinov, Grigory, et al.
Veröffentlicht: (2026)
von: Arshinov, Grigory, et al.
Veröffentlicht: (2026)
On the Spatial Structure of Mixture-of-Experts in Transformers
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
Exploring the Hidden Capacity of LLMs for One-Step Text Generation
von: Mezentsev, Gleb, et al.
Veröffentlicht: (2025)
von: Mezentsev, Gleb, et al.
Veröffentlicht: (2025)
Scaling Laws for Post Training Quantized Large Language Models
von: Xu, Zifei, et al.
Veröffentlicht: (2024)
von: Xu, Zifei, et al.
Veröffentlicht: (2024)
Neural operators meet conjugate gradients: The FCG-NO method for efficient PDE solving
von: Rudikov, Alexander, et al.
Veröffentlicht: (2024)
von: Rudikov, Alexander, et al.
Veröffentlicht: (2024)
The Ky Fan Norms and Beyond: Dual Norms and Combinations for Matrix Optimization
von: Kravatskiy, Alexey, et al.
Veröffentlicht: (2025)
von: Kravatskiy, Alexey, et al.
Veröffentlicht: (2025)
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
von: Laptev, Daniil, et al.
Veröffentlicht: (2025)
von: Laptev, Daniil, et al.
Veröffentlicht: (2025)
Benchmarking Uncertainty Quantification Methods for Large Language Models with LM-Polygraph
von: Vashurin, Roman, et al.
Veröffentlicht: (2024)
von: Vashurin, Roman, et al.
Veröffentlicht: (2024)
Geological Field Restoration through the Lens of Image Inpainting
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2025)
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2025)
Basis Sharing: Cross-Layer Parameter Sharing for Large Language Model Compression
von: Wang, Jingcun, et al.
Veröffentlicht: (2024)
von: Wang, Jingcun, et al.
Veröffentlicht: (2024)
CBQ: Cross-Block Quantization for Large Language Models
von: Ding, Xin, et al.
Veröffentlicht: (2023)
von: Ding, Xin, et al.
Veröffentlicht: (2023)
FBQuant: FeedBack Quantization for Large Language Models
von: Liu, Yijiang, et al.
Veröffentlicht: (2025)
von: Liu, Yijiang, et al.
Veröffentlicht: (2025)
How Quantization Shapes Bias in Large Language Models
von: Marcuzzi, Federico, et al.
Veröffentlicht: (2025)
von: Marcuzzi, Federico, et al.
Veröffentlicht: (2025)
PERELMAN: Pipeline for scientific literature meta-analysis. Technical report
von: Sherki, Daniil, et al.
Veröffentlicht: (2025)
von: Sherki, Daniil, et al.
Veröffentlicht: (2025)
Back to Basics: Revisiting Exploration in Reinforcement Learning for LLM Reasoning via Generative Probabilities
von: Li, Pengyi, et al.
Veröffentlicht: (2026)
von: Li, Pengyi, et al.
Veröffentlicht: (2026)
Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
von: Sivtsov, Danil, et al.
Veröffentlicht: (2025)
von: Sivtsov, Danil, et al.
Veröffentlicht: (2025)
Extreme Compression of Large Language Models via Additive Quantization
von: Egiazarian, Vage, et al.
Veröffentlicht: (2024)
von: Egiazarian, Vage, et al.
Veröffentlicht: (2024)
BAQ: Efficient Bit Allocation Quantization for Large Language Models
von: Zhang, Chao, et al.
Veröffentlicht: (2025)
von: Zhang, Chao, et al.
Veröffentlicht: (2025)
OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
Large Language Models in the Task of Automatic Validation of Text Classifier Predictions
von: Tsymbalov, Aleksandr, et al.
Veröffentlicht: (2025)
von: Tsymbalov, Aleksandr, et al.
Veröffentlicht: (2025)
Another approach to build Lyapunov functions for the first order methods in the quadratic case
von: Merkulov, Daniil, et al.
Veröffentlicht: (2023)
von: Merkulov, Daniil, et al.
Veröffentlicht: (2023)
On the Compressibility of Quantized Large Language Models
von: Mao, Yu, et al.
Veröffentlicht: (2024)
von: Mao, Yu, et al.
Veröffentlicht: (2024)
ApiQ: Finetuning of 2-Bit Quantized Large Language Model
von: Liao, Baohao, et al.
Veröffentlicht: (2024)
von: Liao, Baohao, et al.
Veröffentlicht: (2024)
LCQ: Low-Rank Codebook based Quantization for Large Language Models
von: Cai, Wen-Pu, et al.
Veröffentlicht: (2024)
von: Cai, Wen-Pu, et al.
Veröffentlicht: (2024)
Foundations of Large Language Model Compression -- Part 1: Weight Quantization
von: Young, Sean I.
Veröffentlicht: (2024)
von: Young, Sean I.
Veröffentlicht: (2024)
RSAVQ: Riemannian Sensitivity-Aware Vector Quantization for Large Language Models
von: Xu, Zukang, et al.
Veröffentlicht: (2025)
von: Xu, Zukang, et al.
Veröffentlicht: (2025)
QuIP: 2-Bit Quantization of Large Language Models With Guarantees
von: Chee, Jerry, et al.
Veröffentlicht: (2023)
von: Chee, Jerry, et al.
Veröffentlicht: (2023)
Optimizing Large Language Model Training Using FP4 Quantization
von: Wang, Ruizhe, et al.
Veröffentlicht: (2025)
von: Wang, Ruizhe, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Run LoRA Run: Faster and Lighter LoRA Implementations
von: Cherniuk, Daria, et al.
Veröffentlicht: (2023) -
LoTR: Low Tensor Rank Weight Adaptation
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2024) -
Deep Learning for Subspace Regression
von: Fanaskov, Vladimir, et al.
Veröffentlicht: (2025) -
Bayesian Inverse Problems Meet Flow Matching: Efficient and Flexible Inference via Transformers
von: Sherki, Daniil, et al.
Veröffentlicht: (2025) -
ConDiff: A Challenging Dataset for Neural Solvers of Partial Differential Equations
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2024)