Apertus LLM Family Expansion via Distillation and Quantization
Fuente:
arXiv
Salvato in:
| Autori principali: | Panferov, Andrei, Melikidze, Davit, Jaggi, Martin, Alistarh, Dan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
di: Panferov, Andrei, et al.
Pubblicazione: (2026)
di: Panferov, Andrei, et al.
Pubblicazione: (2026)
CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
di: Tabesh, Soroush, et al.
Pubblicazione: (2025)
di: Tabesh, Soroush, et al.
Pubblicazione: (2025)
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
di: Malinovskii, Vladimir, et al.
Pubblicazione: (2024)
di: Malinovskii, Vladimir, et al.
Pubblicazione: (2024)
Extreme Compression of Large Language Models via Additive Quantization
di: Egiazarian, Vage, et al.
Pubblicazione: (2024)
di: Egiazarian, Vage, et al.
Pubblicazione: (2024)
Grid Games: The Power of Multiple Grids for Quantizing Large Language Models
di: Egiazarian, Vage, et al.
Pubblicazione: (2026)
di: Egiazarian, Vage, et al.
Pubblicazione: (2026)
Homonym Sense Disambiguation in the Georgian Language
di: Melikidze, Davit, et al.
Pubblicazione: (2024)
di: Melikidze, Davit, et al.
Pubblicazione: (2024)
QuEST: Stable Training of LLMs with 1-Bit Weights and Activations
di: Panferov, Andrei, et al.
Pubblicazione: (2025)
di: Panferov, Andrei, et al.
Pubblicazione: (2025)
Unified Scaling Laws for Compressed Representations
di: Panferov, Andrei, et al.
Pubblicazione: (2025)
di: Panferov, Andrei, et al.
Pubblicazione: (2025)
Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
di: Egiazarian, Vage, et al.
Pubblicazione: (2025)
di: Egiazarian, Vage, et al.
Pubblicazione: (2025)
Statistically-Lossless Quantization of Large Language Models
di: Helcig, Michael, et al.
Pubblicazione: (2026)
di: Helcig, Michael, et al.
Pubblicazione: (2026)
WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
di: Chen, Jiale, et al.
Pubblicazione: (2025)
di: Chen, Jiale, et al.
Pubblicazione: (2025)
MatGPTQ: Accurate and Efficient Post-Training Matryoshka Quantization
di: Kleinegger, Maximilian, et al.
Pubblicazione: (2026)
di: Kleinegger, Maximilian, et al.
Pubblicazione: (2026)
Quartet: Native FP4 Training Can Be Optimal for Large Language Models
di: Castro, Roberto L., et al.
Pubblicazione: (2025)
di: Castro, Roberto L., et al.
Pubblicazione: (2025)
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
di: Apertus, Project, et al.
Pubblicazione: (2025)
di: Apertus, Project, et al.
Pubblicazione: (2025)
Correlated Quantization for Faster Nonconvex Distributed Optimization
di: Panferov, Andrei, et al.
Pubblicazione: (2024)
di: Panferov, Andrei, et al.
Pubblicazione: (2024)
Model Compression with Exact Budget Constraints via Riemannian Manifolds
di: Helcig, Michael, et al.
Pubblicazione: (2026)
di: Helcig, Michael, et al.
Pubblicazione: (2026)
Efficient Data Selection at Scale via Influence Distillation
di: Nikdan, Mahdi, et al.
Pubblicazione: (2025)
di: Nikdan, Mahdi, et al.
Pubblicazione: (2025)
The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithm
di: Chen, Jiale, et al.
Pubblicazione: (2025)
di: Chen, Jiale, et al.
Pubblicazione: (2025)
"Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization
di: Kurtic, Eldar, et al.
Pubblicazione: (2024)
di: Kurtic, Eldar, et al.
Pubblicazione: (2024)
Beyond Outliers: A Study of Optimizers Under Quantization
di: Vlassis, Georgios, et al.
Pubblicazione: (2025)
di: Vlassis, Georgios, et al.
Pubblicazione: (2025)
Benchmarking Optimizers for Large Language Model Pretraining
di: Semenov, Andrei, et al.
Pubblicazione: (2025)
di: Semenov, Andrei, et al.
Pubblicazione: (2025)
ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning
di: Melikidze, Davit, et al.
Pubblicazione: (2026)
di: Melikidze, Davit, et al.
Pubblicazione: (2026)
Behemoth: Benchmarking Unlearning in LLMs Using Fully Synthetic Data
di: Iofinova, Eugenia, et al.
Pubblicazione: (2026)
di: Iofinova, Eugenia, et al.
Pubblicazione: (2026)
Compression Scaling Laws:Unifying Sparsity and Quantization
di: Frantar, Elias, et al.
Pubblicazione: (2025)
di: Frantar, Elias, et al.
Pubblicazione: (2025)
ECO: Quantized Training without Full-Precision Master Weights
di: Nikdan, Mahdi, et al.
Pubblicazione: (2026)
di: Nikdan, Mahdi, et al.
Pubblicazione: (2026)
Panza: Design and Analysis of a Fully-Local Personalized Text Writing Assistant
di: Nicolicioiu, Armand, et al.
Pubblicazione: (2024)
di: Nicolicioiu, Armand, et al.
Pubblicazione: (2024)
Gradient-Normalized Smoothness for Optimization with Approximate Hessians
di: Semenov, Andrei, et al.
Pubblicazione: (2025)
di: Semenov, Andrei, et al.
Pubblicazione: (2025)
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
di: Dadgarnia, Alireza, et al.
Pubblicazione: (2026)
di: Dadgarnia, Alireza, et al.
Pubblicazione: (2026)
Layer-wise Quantization for Quantized Optimistic Dual Averaging
di: Nguyen, Anh Duc, et al.
Pubblicazione: (2025)
di: Nguyen, Anh Duc, et al.
Pubblicazione: (2025)
QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
di: Ashkboos, Saleh, et al.
Pubblicazione: (2024)
di: Ashkboos, Saleh, et al.
Pubblicazione: (2024)
EvoPress: Accurate Dynamic Model Compression via Evolutionary Search
di: Sieberling, Oliver, et al.
Pubblicazione: (2024)
di: Sieberling, Oliver, et al.
Pubblicazione: (2024)
Mitigating the Impact of Outlier Channels for Language Model Quantization with Activation Regularization
di: Nrusimha, Aniruddha, et al.
Pubblicazione: (2024)
di: Nrusimha, Aniruddha, et al.
Pubblicazione: (2024)
LLMQ: Efficient Lower-Precision Pretraining for Consumer GPUs
di: Schultheis, Erik, et al.
Pubblicazione: (2025)
di: Schultheis, Erik, et al.
Pubblicazione: (2025)
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
di: Modoranu, Ionut-Vlad, et al.
Pubblicazione: (2026)
di: Modoranu, Ionut-Vlad, et al.
Pubblicazione: (2026)
Powerset Convolutional Neural Networks
di: Wendler, Chris, et al.
Pubblicazione: (2019)
di: Wendler, Chris, et al.
Pubblicazione: (2019)
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
di: Messmer, Bettina, et al.
Pubblicazione: (2025)
di: Messmer, Bettina, et al.
Pubblicazione: (2025)
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models
di: Shutova, Alina, et al.
Pubblicazione: (2025)
di: Shutova, Alina, et al.
Pubblicazione: (2025)
The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
di: Lee, Kwanhee, et al.
Pubblicazione: (2025)
di: Lee, Kwanhee, et al.
Pubblicazione: (2025)
Simple Opinion Dynamics for No-Regret Learning
di: Lazarsfeld, John, et al.
Pubblicazione: (2023)
di: Lazarsfeld, John, et al.
Pubblicazione: (2023)
Towards Fully FP8 GEMM LLM Training at Scale
di: Hernández-Cano, Alejandro, et al.
Pubblicazione: (2025)
di: Hernández-Cano, Alejandro, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
di: Panferov, Andrei, et al.
Pubblicazione: (2026) -
CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
di: Tabesh, Soroush, et al.
Pubblicazione: (2025) -
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
di: Malinovskii, Vladimir, et al.
Pubblicazione: (2024) -
Extreme Compression of Large Language Models via Additive Quantization
di: Egiazarian, Vage, et al.
Pubblicazione: (2024) -
Grid Games: The Power of Multiple Grids for Quantizing Large Language Models
di: Egiazarian, Vage, et al.
Pubblicazione: (2026)