Efficient Pre-Training of LLMs through Truncated SVD Layers
Fuente:
arXiv
Saved in:
| Main Authors: | Kamali, Kaivan, Schweighofer, Kajetan, Shahrzad, Hormoz, Francon, Olivier, Hodjat, Babak, Miikkulainen, Risto |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EVOTER: Evolution of Transparent Explainable Rule-sets
by: Shahrzad, Hormoz, et al.
Published: (2022)
by: Shahrzad, Hormoz, et al.
Published: (2022)
Optimizing the Design of an Artificial Pancreas to Improve Diabetes Management
by: Khanna, Ashok, et al.
Published: (2024)
by: Khanna, Ashok, et al.
Published: (2024)
Overcoming Forgetting in LLM Fine-Tuning with Evolution Strategies
by: Schweighofer, Kajetan, et al.
Published: (2026)
by: Schweighofer, Kajetan, et al.
Published: (2026)
Discovering Effective Policies for Land-Use Planning with Neuroevolution
by: Young, Daniel, et al.
Published: (2023)
by: Young, Daniel, et al.
Published: (2023)
Unlocking the Potential of Global Human Expertise
by: Meyerson, Elliot, et al.
Published: (2024)
by: Meyerson, Elliot, et al.
Published: (2024)
TerraLingua: Emergence and Analysis of Open-endedness in LLM Ecologies
by: Paolo, Giuseppe, et al.
Published: (2026)
by: Paolo, Giuseppe, et al.
Published: (2026)
Solving a Million-Step LLM Task with Zero Errors
by: Meyerson, Elliot, et al.
Published: (2025)
by: Meyerson, Elliot, et al.
Published: (2025)
Asynchronous Evolution of Deep Neural Network Architectures
by: Liang, Jason, et al.
Published: (2023)
by: Liang, Jason, et al.
Published: (2023)
Learning from the Past: How Previous Technological Transformations Can Guide AI Development
by: Miikkulainen, Risto, et al.
Published: (2019)
by: Miikkulainen, Risto, et al.
Published: (2019)
Improving Uncertainty Estimation through Semantically Diverse Language Generation
by: Aichberger, Lukas, et al.
Published: (2024)
by: Aichberger, Lukas, et al.
Published: (2024)
GPU-Accelerated Rule Evaluation and Evolution
by: Shahrzad, Hormoz, et al.
Published: (2024)
by: Shahrzad, Hormoz, et al.
Published: (2024)
Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost
by: Xu, Yinggan, et al.
Published: (2026)
by: Xu, Yinggan, et al.
Published: (2026)
On Information-Theoretic Measures of Predictive Uncertainty
by: Schweighofer, Kajetan, et al.
Published: (2024)
by: Schweighofer, Kajetan, et al.
Published: (2024)
Addressing Pitfalls in the Evaluation of Uncertainty Estimation Methods for Natural Language Generation
by: Ielanskyi, Mykyta, et al.
Published: (2025)
by: Ielanskyi, Mykyta, et al.
Published: (2025)
Leveraging Evolutionary Surrogate-Assisted Prescription in Multi-Objective Chlorination Control Systems
by: Monsia, Rivaaj, et al.
Published: (2025)
by: Monsia, Rivaaj, et al.
Published: (2025)
The Disparate Benefits of Deep Ensembles
by: Schweighofer, Kajetan, et al.
Published: (2024)
by: Schweighofer, Kajetan, et al.
Published: (2024)
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
by: Qiu, Xin, et al.
Published: (2025)
by: Qiu, Xin, et al.
Published: (2025)
The Odyssey of the Fittest: Can Agents Survive and Still Be Good?
by: Waldner, Dylan, et al.
Published: (2025)
by: Waldner, Dylan, et al.
Published: (2025)
Spectral Compact Training: Pre-Training Large Language Models via Permanent Truncated SVD and Stiefel QR Retraction
by: Kohlberger, Björn Roman
Published: (2026)
by: Kohlberger, Björn Roman
Published: (2026)
Optimize Wider, Not Deeper: Consensus Aggregation for Policy Optimization
by: Su, Zelal, et al.
Published: (2026)
by: Su, Zelal, et al.
Published: (2026)
Rethinking Uncertainty Estimation in LLMs: A Principled Single-Sequence Measure
by: Aichberger, Lukas, et al.
Published: (2024)
by: Aichberger, Lukas, et al.
Published: (2024)
DipSVD: Dual-importance Protected SVD for Efficient LLM Compression
by: Ding, Xuan, et al.
Published: (2025)
by: Ding, Xuan, et al.
Published: (2025)
NeuroBack: Improving CDCL SAT Solving using Graph Neural Networks
by: Wang, Wenxi, et al.
Published: (2021)
by: Wang, Wenxi, et al.
Published: (2021)
The Blessing of Dimensionality in LLM Fine-tuning: A Variance-Curvature Perspective
by: Liang, Qiyao, et al.
Published: (2026)
by: Liang, Qiyao, et al.
Published: (2026)
Semantic Density: Uncertainty Quantification for Large Language Models through Confidence Measurement in Semantic Space
by: Qiu, Xin, et al.
Published: (2024)
by: Qiu, Xin, et al.
Published: (2024)
CoLA: Compute-Efficient Pre-Training of LLMs via Low-Rank Activation
by: Liu, Ziyue, et al.
Published: (2025)
by: Liu, Ziyue, et al.
Published: (2025)
SAFE-SVD: Sensitivity-Aware Fidelity-Enforcing SVD for Physics Foundation Models
by: Hong, Chengjie, et al.
Published: (2026)
by: Hong, Chengjie, et al.
Published: (2026)
Layer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-Training
by: Nepal, Aadim, et al.
Published: (2025)
by: Nepal, Aadim, et al.
Published: (2025)
Efficient Knowledge Deletion from Trained Models through Layer-wise Partial Machine Unlearning
by: Gogineni, Vinay Chakravarthi, et al.
Published: (2024)
by: Gogineni, Vinay Chakravarthi, et al.
Published: (2024)
Pre-Training LLMs on a budget: A comparison of three optimizers
by: Schlotthauer, Joel, et al.
Published: (2025)
by: Schlotthauer, Joel, et al.
Published: (2025)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
by: Shao, Zishan, et al.
Published: (2025)
by: Shao, Zishan, et al.
Published: (2025)
Neural Cellular Automata for ARC-AGI
by: Xu, Kevin, et al.
Published: (2025)
by: Xu, Kevin, et al.
Published: (2025)
Enhancing Delta Compression in LLMs via SVD-based Quantization Error Minimization
by: Xiong, Boya, et al.
Published: (2025)
by: Xiong, Boya, et al.
Published: (2025)
TEON: Tensorized Orthonormalization Beyond Layer-Wise Muon for Large Language Model Pre-Training
by: Zhang, Ruijie, et al.
Published: (2026)
by: Zhang, Ruijie, et al.
Published: (2026)
The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation
by: Lan, Yifan, et al.
Published: (2026)
by: Lan, Yifan, et al.
Published: (2026)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
by: Zhao, Zhixiong, et al.
Published: (2025)
by: Zhao, Zhixiong, et al.
Published: (2025)
Post-processing for Fair Regression via Explainable SVD
by: Zuo, Zhiqun, et al.
Published: (2025)
by: Zuo, Zhiqun, et al.
Published: (2025)
Enhancing Pre-Trained Model-Based Class-Incremental Learning through Neural Collapse
by: He, Kun, et al.
Published: (2025)
by: He, Kun, et al.
Published: (2025)
Evolution With Purpose: Hierarchy-Informed Optimization of Whole-Brain Models
by: Shahrzad, Hormoz, et al.
Published: (2026)
by: Shahrzad, Hormoz, et al.
Published: (2026)
Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination
by: Liang, Qiyao, et al.
Published: (2026)
by: Liang, Qiyao, et al.
Published: (2026)
Similar Items
-
EVOTER: Evolution of Transparent Explainable Rule-sets
by: Shahrzad, Hormoz, et al.
Published: (2022) -
Optimizing the Design of an Artificial Pancreas to Improve Diabetes Management
by: Khanna, Ashok, et al.
Published: (2024) -
Overcoming Forgetting in LLM Fine-Tuning with Evolution Strategies
by: Schweighofer, Kajetan, et al.
Published: (2026) -
Discovering Effective Policies for Land-Use Planning with Neuroevolution
by: Young, Daniel, et al.
Published: (2023) -
Unlocking the Potential of Global Human Expertise
by: Meyerson, Elliot, et al.
Published: (2024)