MetaTT: A Global Tensor-Train Adapter for Parameter-Efficient Fine-Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lopez-Piqueres, Javier, Deshpande, Pranav, Ray, Archan, Villani, Mattia J., Pistoia, Marco, Kumar, Niraj
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909902954299392
author Lopez-Piqueres, Javier
Deshpande, Pranav
Ray, Archan
Villani, Mattia J.
Pistoia, Marco
Kumar, Niraj
author_facet Lopez-Piqueres, Javier
Deshpande, Pranav
Ray, Archan
Villani, Mattia J.
Pistoia, Marco
Kumar, Niraj
contents We present MetaTT, a Tensor Train (TT) adapter framework for fine-tuning of pre-trained transformers. MetaTT enables flexible and parameter-efficient model adaptation by using a single shared TT to factorize transformer sub-modules. This factorization indexes key structural dimensions, including layer and matrix type, and can optionally incorporate heads and tasks. This design allows MetaTT's parameter count to scale with the sum, rather than the product, of the modes, resulting in a substantially more compact adapter. Our benchmarks compare MetaTT with LoRA along with recent state-of-the-art matrix and tensor decomposition based fine-tuning methods. We observe that when tested on single-task standard language modeling benchmarks, MetaTT achieves competitive parameter efficiency to accuracy tradeoff. We further demonstrate that MetaTT performs competitively when compared to state-of-the-art methods on multi-task learning. Finally, we leverage the TT-ansatz to design a rank adaptive optimizer inspired by the DMRG method from many-body physics. Our results demonstrate that integrating this approach with AdamW enhances optimization performance for a specified target rank.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09105
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MetaTT: A Global Tensor-Train Adapter for Parameter-Efficient Fine-Tuning
Lopez-Piqueres, Javier
Deshpande, Pranav
Ray, Archan
Villani, Mattia J.
Pistoia, Marco
Kumar, Niraj
Machine Learning
Artificial Intelligence
Quantum Physics
We present MetaTT, a Tensor Train (TT) adapter framework for fine-tuning of pre-trained transformers. MetaTT enables flexible and parameter-efficient model adaptation by using a single shared TT to factorize transformer sub-modules. This factorization indexes key structural dimensions, including layer and matrix type, and can optionally incorporate heads and tasks. This design allows MetaTT's parameter count to scale with the sum, rather than the product, of the modes, resulting in a substantially more compact adapter. Our benchmarks compare MetaTT with LoRA along with recent state-of-the-art matrix and tensor decomposition based fine-tuning methods. We observe that when tested on single-task standard language modeling benchmarks, MetaTT achieves competitive parameter efficiency to accuracy tradeoff. We further demonstrate that MetaTT performs competitively when compared to state-of-the-art methods on multi-task learning. Finally, we leverage the TT-ansatz to design a rank adaptive optimizer inspired by the DMRG method from many-body physics. Our results demonstrate that integrating this approach with AdamW enhances optimization performance for a specified target rank.
title MetaTT: A Global Tensor-Train Adapter for Parameter-Efficient Fine-Tuning
topic Machine Learning
Artificial Intelligence
Quantum Physics
url https://arxiv.org/abs/2506.09105