Gespeichert in:
| Hauptverfasser: | Volkova, Alexandra, Safaryan, Mher, Lampert, Christoph H., Alistarh, Dan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.07712 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unified Scaling Laws for Compressed Representations
von: Panferov, Andrei, et al.
Veröffentlicht: (2025)
von: Panferov, Andrei, et al.
Veröffentlicht: (2025)
CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
von: Tabesh, Soroush, et al.
Veröffentlicht: (2025)
von: Tabesh, Soroush, et al.
Veröffentlicht: (2025)
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)
LDAdam: Adaptive Optimization from Low-Dimensional Gradient Statistics
von: Robert, Thomas, et al.
Veröffentlicht: (2024)
von: Robert, Thomas, et al.
Veröffentlicht: (2024)
DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)
The Iterative Optimal Brain Surgeon: Faster Sparse Recovery by Leveraging Second-Order Information
von: Wu, Diyuan, et al.
Veröffentlicht: (2024)
von: Wu, Diyuan, et al.
Veröffentlicht: (2024)
FFT-based Dynamic Subspace Selection for Low-Rank Adaptive Optimization of Large Language Models
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2025)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2025)
MicroAdam: Accurate Adaptive Optimization with Low Space Overhead and Provable Convergence
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2024)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2024)
Beyond Outliers: A Study of Optimizers Under Quantization
von: Vlassis, Georgios, et al.
Veröffentlicht: (2025)
von: Vlassis, Georgios, et al.
Veröffentlicht: (2025)
GradSkip: Communication-Accelerated Local Gradient Methods with Better Computational Complexity
von: Maranjyan, Artavazd, et al.
Veröffentlicht: (2022)
von: Maranjyan, Artavazd, et al.
Veröffentlicht: (2022)
LoRDO: Distributed Low-Rank Optimization with Infrequent Communication
von: Jovanović, Andrej, et al.
Veröffentlicht: (2026)
von: Jovanović, Andrej, et al.
Veröffentlicht: (2026)
On Biased Compression for Distributed Learning
von: Beznosikov, Aleksandr, et al.
Veröffentlicht: (2020)
von: Beznosikov, Aleksandr, et al.
Veröffentlicht: (2020)
Compression Scaling Laws:Unifying Sparsity and Quantization
von: Frantar, Elias, et al.
Veröffentlicht: (2025)
von: Frantar, Elias, et al.
Veröffentlicht: (2025)
MT-DAO: Multi-Timescale Distributed Adaptive Optimizers with Local Updates
von: Iacob, Alex, et al.
Veröffentlicht: (2025)
von: Iacob, Alex, et al.
Veröffentlicht: (2025)
Behemoth: Benchmarking Unlearning in LLMs Using Fully Synthetic Data
von: Iofinova, Eugenia, et al.
Veröffentlicht: (2026)
von: Iofinova, Eugenia, et al.
Veröffentlicht: (2026)
Model Compression with Exact Budget Constraints via Riemannian Manifolds
von: Helcig, Michael, et al.
Veröffentlicht: (2026)
von: Helcig, Michael, et al.
Veröffentlicht: (2026)
Towards Combinatorial Interpretability of Neural Computation
von: Adler, Micah, et al.
Veröffentlicht: (2025)
von: Adler, Micah, et al.
Veröffentlicht: (2025)
DES-LOC: Desynced Low Communication Adaptive Optimizers for Training Foundation Models
von: Iacob, Alex, et al.
Veröffentlicht: (2025)
von: Iacob, Alex, et al.
Veröffentlicht: (2025)
ASIDE: Architectural Separation of Instructions and Data in Language Models
von: Zverev, Egor, et al.
Veröffentlicht: (2025)
von: Zverev, Egor, et al.
Veröffentlicht: (2025)
LLMQ: Efficient Lower-Precision Pretraining for Consumer GPUs
von: Schultheis, Erik, et al.
Veröffentlicht: (2025)
von: Schultheis, Erik, et al.
Veröffentlicht: (2025)
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws
von: Jin, Tian, et al.
Veröffentlicht: (2025)
von: Jin, Tian, et al.
Veröffentlicht: (2025)
Efficient Data Selection at Scale via Influence Distillation
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2025)
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2025)
Hybrid Decentralized Optimization: Leveraging Both First- and Zeroth-Order Optimizers for Faster Convergence
von: Ansaripour, Matin, et al.
Veröffentlicht: (2022)
von: Ansaripour, Matin, et al.
Veröffentlicht: (2022)
MatGPTQ: Accurate and Efficient Post-Training Matryoshka Quantization
von: Kleinegger, Maximilian, et al.
Veröffentlicht: (2026)
von: Kleinegger, Maximilian, et al.
Veröffentlicht: (2026)
Statistically-Lossless Quantization of Large Language Models
von: Helcig, Michael, et al.
Veröffentlicht: (2026)
von: Helcig, Michael, et al.
Veröffentlicht: (2026)
Powerset Convolutional Neural Networks
von: Wendler, Chris, et al.
Veröffentlicht: (2019)
von: Wendler, Chris, et al.
Veröffentlicht: (2019)
RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2024)
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2024)
Simple Opinion Dynamics for No-Regret Learning
von: Lazarsfeld, John, et al.
Veröffentlicht: (2023)
von: Lazarsfeld, John, et al.
Veröffentlicht: (2023)
HALO: Hadamard-Assisted Lower-Precision Optimization for LLMs
von: Ashkboos, Saleh, et al.
Veröffentlicht: (2025)
von: Ashkboos, Saleh, et al.
Veröffentlicht: (2025)
Apertus LLM Family Expansion via Distillation and Quantization
von: Panferov, Andrei, et al.
Veröffentlicht: (2026)
von: Panferov, Andrei, et al.
Veröffentlicht: (2026)
Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
von: Panferov, Andrei, et al.
Veröffentlicht: (2026)
von: Panferov, Andrei, et al.
Veröffentlicht: (2026)
EvoPress: Accurate Dynamic Model Compression via Evolutionary Search
von: Sieberling, Oliver, et al.
Veröffentlicht: (2024)
von: Sieberling, Oliver, et al.
Veröffentlicht: (2024)
Communication-Efficient Federated Learning With Data and Client Heterogeneity
von: Zakerinia, Hossein, et al.
Veröffentlicht: (2022)
von: Zakerinia, Hossein, et al.
Veröffentlicht: (2022)
Towards Scaling Laws for Symbolic Regression
von: Otte, David, et al.
Veröffentlicht: (2025)
von: Otte, David, et al.
Veröffentlicht: (2025)
Adaptive Sampling and Clipping for Private Worst-Case Group Optimization
von: Cairney-Leeming, Max, et al.
Veröffentlicht: (2026)
von: Cairney-Leeming, Max, et al.
Veröffentlicht: (2026)
SPADE: Sparsity-Guided Debugging for Deep Neural Networks
von: Moakhar, Arshia Soltani, et al.
Veröffentlicht: (2023)
von: Moakhar, Arshia Soltani, et al.
Veröffentlicht: (2023)
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
von: Kurtic, Eldar, et al.
Veröffentlicht: (2024)
von: Kurtic, Eldar, et al.
Veröffentlicht: (2024)
From Privacy to Generalization: Linear Max-Information Bounds for DP-SGD
von: Lampert, Christoph H., et al.
Veröffentlicht: (2026)
von: Lampert, Christoph H., et al.
Veröffentlicht: (2026)
Fast Rate Bounds for Multi-Task and Meta-Learning with Different Sample Sizes
von: Zakerinia, Hossein, et al.
Veröffentlicht: (2025)
von: Zakerinia, Hossein, et al.
Veröffentlicht: (2025)
1-Lipschitz Neural Networks are more expressive with N-Activations
von: Prach, Bernd, et al.
Veröffentlicht: (2023)
von: Prach, Bernd, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Unified Scaling Laws for Compressed Representations
von: Panferov, Andrei, et al.
Veröffentlicht: (2025) -
CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
von: Tabesh, Soroush, et al.
Veröffentlicht: (2025) -
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026) -
LDAdam: Adaptive Optimization from Low-Dimensional Gradient Statistics
von: Robert, Thomas, et al.
Veröffentlicht: (2024) -
DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)