CoLA: Compute-Efficient Pre-Training of LLMs via Low-Rank Activation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Ziyue, Zhang, Ruijie, Wang, Zhengyang, Yan, Mingsong, Yang, Zi, Hovland, Paul, Nicolae, Bogdan, Cappello, Franck, Tang, Sui, Zhang, Zheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
von: Wang, Zhengyang, et al.
Veröffentlicht: (2025)
von: Wang, Zhengyang, et al.
Veröffentlicht: (2025)
ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload
von: Liu, Ziyue, et al.
Veröffentlicht: (2026)
von: Liu, Ziyue, et al.
Veröffentlicht: (2026)
LaX: Boosting Low-Rank Training of Foundation Models via Latent Crossing
von: Zhang, Ruijie, et al.
Veröffentlicht: (2025)
von: Zhang, Ruijie, et al.
Veröffentlicht: (2025)
CoLA: Collaborative Low-Rank Adaptation
von: Zhou, Yiyun, et al.
Veröffentlicht: (2025)
von: Zhou, Yiyun, et al.
Veröffentlicht: (2025)
MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall
von: Maurya, Avinash, et al.
Veröffentlicht: (2025)
von: Maurya, Avinash, et al.
Veröffentlicht: (2025)
CoMERA: Computing- and Memory-Efficient Training via Rank-Adaptive Tensor Optimization
von: Yang, Zi, et al.
Veröffentlicht: (2024)
von: Yang, Zi, et al.
Veröffentlicht: (2024)
Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
TEON: Tensorized Orthonormalization Beyond Layer-Wise Muon for Large Language Model Pre-Training
von: Zhang, Ruijie, et al.
Veröffentlicht: (2026)
von: Zhang, Ruijie, et al.
Veröffentlicht: (2026)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
MUON+: Towards More Effective Muon via One Additional Normalization Step for LLM Pre-training
von: Zhang, Ruijie, et al.
Veröffentlicht: (2026)
von: Zhang, Ruijie, et al.
Veröffentlicht: (2026)
Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning
von: Liu, Ziyue, et al.
Veröffentlicht: (2026)
von: Liu, Ziyue, et al.
Veröffentlicht: (2026)
CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks
von: Suharitdamrong, Wish, et al.
Veröffentlicht: (2026)
von: Suharitdamrong, Wish, et al.
Veröffentlicht: (2026)
SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs
von: Lee, Jin, et al.
Veröffentlicht: (2026)
von: Lee, Jin, et al.
Veröffentlicht: (2026)
CoLA: A Choice Leakage Attack Framework to Expose Privacy Risks in Subset Training
von: Li, Qi, et al.
Veröffentlicht: (2026)
von: Li, Qi, et al.
Veröffentlicht: (2026)
CoLA: Conditional Dropout and Language-driven Robust Dual-modal Salient Object Detection
von: Hao, Shuang, et al.
Veröffentlicht: (2024)
von: Hao, Shuang, et al.
Veröffentlicht: (2024)
CyclicFL: A Cyclic Model Pre-Training Approach to Efficient Federated Learning
von: Zhang, Pengyu, et al.
Veröffentlicht: (2023)
von: Zhang, Pengyu, et al.
Veröffentlicht: (2023)
On the Convergence and Size Transferability of Continuous-depth Graph Neural Networks
von: Yan, Mingsong, et al.
Veröffentlicht: (2025)
von: Yan, Mingsong, et al.
Veröffentlicht: (2025)
MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization
von: Su, Yupeng, et al.
Veröffentlicht: (2026)
von: Su, Yupeng, et al.
Veröffentlicht: (2026)
ELAS: Efficient Pre-Training of Low-Rank Large Language Models via 2:4 Activation Sparsity
von: Li, Jiaxi, et al.
Veröffentlicht: (2026)
von: Li, Jiaxi, et al.
Veröffentlicht: (2026)
Collaborative Low-Rank Adaptation for Pre-Trained Vision Transformers
von: Liu, Zheng, et al.
Veröffentlicht: (2025)
von: Liu, Zheng, et al.
Veröffentlicht: (2025)
Hardware-Software Co-design for Distributed Quantum Computing
von: Liu, Ji, et al.
Veröffentlicht: (2025)
von: Liu, Ji, et al.
Veröffentlicht: (2025)
Heterogeneous Low-Bandwidth Pre-Training of LLMs
von: Obeidi, Yazan, et al.
Veröffentlicht: (2026)
von: Obeidi, Yazan, et al.
Veröffentlicht: (2026)
LASER: Low-Rank Activation SVD for Efficient Recursion
von: Çakar, Ege, et al.
Veröffentlicht: (2026)
von: Çakar, Ege, et al.
Veröffentlicht: (2026)
ELRT: Efficient Low-Rank Training for Compact Convolutional Neural Networks
von: Sui, Yang, et al.
Veröffentlicht: (2024)
von: Sui, Yang, et al.
Veröffentlicht: (2024)
VSS Challenge Problem: Verifying the Correctness of AllReduce Algorithms in the MPICH Implementation of MPI
von: Hovland, Paul D.
Veröffentlicht: (2025)
von: Hovland, Paul D.
Veröffentlicht: (2025)
PowLU: An Activation Function for Stable Pre-Training of LLMs
von: Jiang, Peijie, et al.
Veröffentlicht: (2026)
von: Jiang, Peijie, et al.
Veröffentlicht: (2026)
Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM
von: Solgi, Ryan, et al.
Veröffentlicht: (2025)
von: Solgi, Ryan, et al.
Veröffentlicht: (2025)
Breaking the Blocks: Continuous Low-Rank Decomposed Scaling for Unified LLM Quantization and Adaptation
von: Tang, Pingzhi, et al.
Veröffentlicht: (2026)
von: Tang, Pingzhi, et al.
Veröffentlicht: (2026)
LANCE: Low Rank Activation Compression for Efficient On-Device Continual Learning
von: Apolinario, Marco Paul E., et al.
Veröffentlicht: (2025)
von: Apolinario, Marco Paul E., et al.
Veröffentlicht: (2025)
Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression
von: Yan, Mingsong, et al.
Veröffentlicht: (2026)
von: Yan, Mingsong, et al.
Veröffentlicht: (2026)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
von: Arif, Moiz, et al.
Veröffentlicht: (2026)
von: Arif, Moiz, et al.
Veröffentlicht: (2026)
FedSLoP: Memory-Efficient Federated Learning with Low-Rank Gradient Projection
von: He, Yutong, et al.
Veröffentlicht: (2026)
von: He, Yutong, et al.
Veröffentlicht: (2026)
Constructing Level Sets Using Smoothed Approximate Bayesian Computation
von: Edwards, David, et al.
Veröffentlicht: (2024)
von: Edwards, David, et al.
Veröffentlicht: (2024)
FuRA: Full-Rank Parameter-Efficient Fine-Tuning with Spectral Preconditioning
von: Zhao, Yequan, et al.
Veröffentlicht: (2026)
von: Zhao, Yequan, et al.
Veröffentlicht: (2026)
The Restriction of Efficient Geodesics to the Non-separating Complex of Curves
von: Hovland, Seth, et al.
Veröffentlicht: (2023)
von: Hovland, Seth, et al.
Veröffentlicht: (2023)
Differentiating Through Linear Solvers
von: Hovland, Paul, et al.
Veröffentlicht: (2024)
von: Hovland, Paul, et al.
Veröffentlicht: (2024)
Low-Rank Quantization-Aware Training for LLMs
von: Bondarenko, Yelysei, et al.
Veröffentlicht: (2024)
von: Bondarenko, Yelysei, et al.
Veröffentlicht: (2024)
Understanding The Effectiveness of Lossy Compression in Machine Learning Training Sets
von: Underwood, Robert, et al.
Veröffentlicht: (2024)
von: Underwood, Robert, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
von: Wang, Zhengyang, et al.
Veröffentlicht: (2025) -
ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload
von: Liu, Ziyue, et al.
Veröffentlicht: (2026) -
LaX: Boosting Low-Rank Training of Foundation Models via Latent Crossing
von: Zhang, Ruijie, et al.
Veröffentlicht: (2025) -
CoLA: Collaborative Low-Rank Adaptation
von: Zhou, Yiyun, et al.
Veröffentlicht: (2025) -
MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall
von: Maurya, Avinash, et al.
Veröffentlicht: (2025)