Text-to-LoRA: Instant Transformer Adaption

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Charakorn, Rujikorn, Cetin, Edoardo, Tang, Yujin, Lange, Robert Tjarko
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909642348560384
author Charakorn, Rujikorn
Cetin, Edoardo
Tang, Yujin
Lange, Robert Tjarko
author_facet Charakorn, Rujikorn
Cetin, Edoardo
Tang, Yujin
Lange, Robert Tjarko
contents While Foundation Models provide a general tool for rapid content creation, they regularly require task-specific adaptation. Traditionally, this exercise involves careful curation of datasets and repeated fine-tuning of the underlying model. Fine-tuning techniques enable practitioners to adapt foundation models for many new applications but require expensive and lengthy training while being notably sensitive to hyperparameter choices. To overcome these limitations, we introduce Text-to-LoRA (T2L), a model capable of adapting large language models (LLMs) on the fly solely based on a natural language description of the target task. T2L is a hypernetwork trained to construct LoRAs in a single inexpensive forward pass. After training T2L on a suite of 9 pre-trained LoRA adapters (GSM8K, Arc, etc.), we show that the ad-hoc reconstructed LoRA instances match the performance of task-specific adapters across the corresponding test sets. Furthermore, T2L can compress hundreds of LoRA instances and zero-shot generalize to entirely unseen tasks. This approach provides a significant step towards democratizing the specialization of foundation models and enables language-based adaptation with minimal compute requirements. Our code is available at https://github.com/SakanaAI/text-to-lora
format Preprint
id arxiv_https___arxiv_org_abs_2506_06105
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Text-to-LoRA: Instant Transformer Adaption
Charakorn, Rujikorn
Cetin, Edoardo
Tang, Yujin
Lange, Robert Tjarko
Machine Learning
Artificial Intelligence
While Foundation Models provide a general tool for rapid content creation, they regularly require task-specific adaptation. Traditionally, this exercise involves careful curation of datasets and repeated fine-tuning of the underlying model. Fine-tuning techniques enable practitioners to adapt foundation models for many new applications but require expensive and lengthy training while being notably sensitive to hyperparameter choices. To overcome these limitations, we introduce Text-to-LoRA (T2L), a model capable of adapting large language models (LLMs) on the fly solely based on a natural language description of the target task. T2L is a hypernetwork trained to construct LoRAs in a single inexpensive forward pass. After training T2L on a suite of 9 pre-trained LoRA adapters (GSM8K, Arc, etc.), we show that the ad-hoc reconstructed LoRA instances match the performance of task-specific adapters across the corresponding test sets. Furthermore, T2L can compress hundreds of LoRA instances and zero-shot generalize to entirely unseen tasks. This approach provides a significant step towards democratizing the specialization of foundation models and enables language-based adaptation with minimal compute requirements. Our code is available at https://github.com/SakanaAI/text-to-lora
title Text-to-LoRA: Instant Transformer Adaption
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.06105