LaX: Boosting Low-Rank Training of Foundation Models via Latent Crossing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Ruijie, Liu, Ziyue, Wang, Zhengyang, Zhang, Zheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909868981485568
author Zhang, Ruijie
Liu, Ziyue
Wang, Zhengyang
Zhang, Zheng
author_facet Zhang, Ruijie
Liu, Ziyue
Wang, Zhengyang
Zhang, Zheng
contents Training foundation models such as ViTs and LLMs requires tremendous computing cost. Low-rank matrix or tensor factorization offers a parameter-efficient alternative, but often downgrades performance due to the restricted parameter space. In this work, we introduce {\textbf{Latent Crossing (LaX)}} -- a simple yet effective plug-and-play module that enhances the capacity of low-rank models by enabling information flow across low-rank subspaces. We extensively validate the benefits of LaX on pre-training tasks with ViT-Base/Large and LLaMA-like models ranging from 60M to 1B parameters. LaX boosts low-rank model performance to match or exceed the full-rank baselines while using 2-3\(\times\) fewer parameters. When equipped with low-rank adapters (i.e., LoRA) for fine-tuning LLaMA-7/13B, LaX consistently improves performance on arithmetic and common sense reasoning tasks with negligible cost.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21732
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LaX: Boosting Low-Rank Training of Foundation Models via Latent Crossing
Zhang, Ruijie
Liu, Ziyue
Wang, Zhengyang
Zhang, Zheng
Machine Learning
Training foundation models such as ViTs and LLMs requires tremendous computing cost. Low-rank matrix or tensor factorization offers a parameter-efficient alternative, but often downgrades performance due to the restricted parameter space. In this work, we introduce {\textbf{Latent Crossing (LaX)}} -- a simple yet effective plug-and-play module that enhances the capacity of low-rank models by enabling information flow across low-rank subspaces. We extensively validate the benefits of LaX on pre-training tasks with ViT-Base/Large and LLaMA-like models ranging from 60M to 1B parameters. LaX boosts low-rank model performance to match or exceed the full-rank baselines while using 2-3\(\times\) fewer parameters. When equipped with low-rank adapters (i.e., LoRA) for fine-tuning LLaMA-7/13B, LaX consistently improves performance on arithmetic and common sense reasoning tasks with negligible cost.
title LaX: Boosting Low-Rank Training of Foundation Models via Latent Crossing
topic Machine Learning
url https://arxiv.org/abs/2505.21732