Improving Recursive Transformers with Mixture of LoRAs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nouriborji, Mohammadmahdi, Rohanian, Morteza, Rohanian, Omid
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912771233284096
author Nouriborji, Mohammadmahdi
Rohanian, Morteza
Rohanian, Omid
author_facet Nouriborji, Mohammadmahdi
Rohanian, Morteza
Rohanian, Omid
contents Parameter sharing in recursive transformers reduces model size but collapses layer-wise expressivity. We propose Mixture of LoRAs (MoL), a lightweight conditional-computation mechanism that inserts Low-Rank Adaptation (LoRA) experts inside a shared feed-forward network (FFN). MoL enables token-conditional weight-space modulation of the shared FFN without untying backbone parameters, unlike prior approaches that add fixed or externally attached adapters. We pretrain a modernised recursive architecture, ModernALBERT, integrating rotary embeddings, GeGLU, FlashAttention, and a distillation-based initialisation. Across GLUE, SQuAD-v2, and BEIR, ModernALBERT (50M--120M) achieves state-of-the-art performance among compact models and surpasses larger fully parameterised baselines. We also propose an expert-merging procedure that compresses MoL into a single adapter at inference while preserving accuracy, enabling efficient deployment. Our results show that conditional weight-space modulation effectively restores the expressivity lost under aggressive parameter sharing in recursive transformers.
format Preprint
id arxiv_https___arxiv_org_abs_2512_12880
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Recursive Transformers with Mixture of LoRAs
Nouriborji, Mohammadmahdi
Rohanian, Morteza
Rohanian, Omid
Machine Learning
68T05, 68T50
Parameter sharing in recursive transformers reduces model size but collapses layer-wise expressivity. We propose Mixture of LoRAs (MoL), a lightweight conditional-computation mechanism that inserts Low-Rank Adaptation (LoRA) experts inside a shared feed-forward network (FFN). MoL enables token-conditional weight-space modulation of the shared FFN without untying backbone parameters, unlike prior approaches that add fixed or externally attached adapters. We pretrain a modernised recursive architecture, ModernALBERT, integrating rotary embeddings, GeGLU, FlashAttention, and a distillation-based initialisation. Across GLUE, SQuAD-v2, and BEIR, ModernALBERT (50M--120M) achieves state-of-the-art performance among compact models and surpasses larger fully parameterised baselines. We also propose an expert-merging procedure that compresses MoL into a single adapter at inference while preserving accuracy, enabling efficient deployment. Our results show that conditional weight-space modulation effectively restores the expressivity lost under aggressive parameter sharing in recursive transformers.
title Improving Recursive Transformers with Mixture of LoRAs
topic Machine Learning
68T05, 68T50
url https://arxiv.org/abs/2512.12880