MoKA: Mixture of Kronecker Adapters

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sadeghi, Mohammadreza, Nejad, Mahsa Ghazvini, Asl, MirHamed Jafarzadeh, Gu, Yu, Yu, Yuanhao, Asgharian, Masoud, Nia, Vahid Partovi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909723664580608
author Sadeghi, Mohammadreza
Nejad, Mahsa Ghazvini
Asl, MirHamed Jafarzadeh
Gu, Yu
Yu, Yuanhao
Asgharian, Masoud
Nia, Vahid Partovi
author_facet Sadeghi, Mohammadreza
Nejad, Mahsa Ghazvini
Asl, MirHamed Jafarzadeh
Gu, Yu
Yu, Yuanhao
Asgharian, Masoud
Nia, Vahid Partovi
contents Parameter-efficient fine-tuning (PEFT) is essential for reducing the computational overhead of large language models (LLMs). Low-rank family adapters are commonly used to control the parameter size efficiently while maintaining the generative power of LLMs. However, their limited expressiveness due to the rank constraint often restricts their performance on complex tasks. We propose Mixture of Kronecker Adapters (MoKA), a new generation of Kronecker adapters that addresses this limitation by modeling weight updates as a mixture of Kronecker products. Our proposed adapter leverages a gating mechanism that measures the importance of each Kronecker factor, enabling more expressive adaptation. Moreover, MoKA enables a rank flexibility that provides a better trade-off between parameter efficiency and accuracy. To ensure hardware efficiency, we reformulate Kronecker computations using standard matrix operations, allowing seamless deployment on GPU-optimized hardware. We conduct extensive experiments on instruction-tuning and commonsense reasoning tasks using low-bit quantized versions of LLaMA2-7B and LLaMA3-8B models. MoKA not only outperforms PEFT baselines, but also reduces the number of trainable parameters up to 27x, achieving state-of-the-art trade-offs between performance and parameter efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03527
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MoKA: Mixture of Kronecker Adapters
Sadeghi, Mohammadreza
Nejad, Mahsa Ghazvini
Asl, MirHamed Jafarzadeh
Gu, Yu
Yu, Yuanhao
Asgharian, Masoud
Nia, Vahid Partovi
Machine Learning
Artificial Intelligence
Computation and Language
Parameter-efficient fine-tuning (PEFT) is essential for reducing the computational overhead of large language models (LLMs). Low-rank family adapters are commonly used to control the parameter size efficiently while maintaining the generative power of LLMs. However, their limited expressiveness due to the rank constraint often restricts their performance on complex tasks. We propose Mixture of Kronecker Adapters (MoKA), a new generation of Kronecker adapters that addresses this limitation by modeling weight updates as a mixture of Kronecker products. Our proposed adapter leverages a gating mechanism that measures the importance of each Kronecker factor, enabling more expressive adaptation. Moreover, MoKA enables a rank flexibility that provides a better trade-off between parameter efficiency and accuracy. To ensure hardware efficiency, we reformulate Kronecker computations using standard matrix operations, allowing seamless deployment on GPU-optimized hardware. We conduct extensive experiments on instruction-tuning and commonsense reasoning tasks using low-bit quantized versions of LLaMA2-7B and LLaMA3-8B models. MoKA not only outperforms PEFT baselines, but also reduces the number of trainable parameters up to 27x, achieving state-of-the-art trade-offs between performance and parameter efficiency.
title MoKA: Mixture of Kronecker Adapters
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.03527