PHLoRA: data-free Post-hoc Low-Rank Adapter extraction from full-rank checkpoint

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vasani, Bhoomit, FitzGerald, Jack, Fang, Anjie, Vaish, Sushmit
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911153034100736
author Vasani, Bhoomit
FitzGerald, Jack
Fang, Anjie
Vaish, Sushmit
author_facet Vasani, Bhoomit
FitzGerald, Jack
Fang, Anjie
Vaish, Sushmit
contents We introduce PHLoRA (Pronounced "flora"). (Post-hoc LoRA), a simple yet powerful method to extract low-rank adaptation adapters from full-rank fine-tuned models without requiring access to training data or gradients. By computing the low-rank decomposition of weight differences between a base model and its fine-tuned counterpart, our method reconstructs adapter modules that can be merged or dynamically routed at inference time via S-LoRA, or served in scalable, industry settings using platforms like NVIDIA NIM. This approach amortizes latency overhead across requests and yields substantial cost savings. Unlike prior work that trains each adapter explicitly, our approach decouples fine-tuning from adapter generation, allowing adapter extraction from existing full-rank models or third-party checkpoints. Experiments on text, image, and video benchmarks using the Amazon Nova model family demonstrate that extracted adapters preserve high energy from the full weight delta, can be pruned safely, and yield negligible degradation in downstream task performance when re-merged. Overall, PHLoRA provides a practical path for making all existing full-rank checkpoints adapter-ready, democratizing scalable inference for all models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_10971
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PHLoRA: data-free Post-hoc Low-Rank Adapter extraction from full-rank checkpoint
Vasani, Bhoomit
FitzGerald, Jack
Fang, Anjie
Vaish, Sushmit
Machine Learning
Artificial Intelligence
We introduce PHLoRA (Pronounced "flora"). (Post-hoc LoRA), a simple yet powerful method to extract low-rank adaptation adapters from full-rank fine-tuned models without requiring access to training data or gradients. By computing the low-rank decomposition of weight differences between a base model and its fine-tuned counterpart, our method reconstructs adapter modules that can be merged or dynamically routed at inference time via S-LoRA, or served in scalable, industry settings using platforms like NVIDIA NIM. This approach amortizes latency overhead across requests and yields substantial cost savings. Unlike prior work that trains each adapter explicitly, our approach decouples fine-tuning from adapter generation, allowing adapter extraction from existing full-rank models or third-party checkpoints. Experiments on text, image, and video benchmarks using the Amazon Nova model family demonstrate that extracted adapters preserve high energy from the full weight delta, can be pruned safely, and yield negligible degradation in downstream task performance when re-merged. Overall, PHLoRA provides a practical path for making all existing full-rank checkpoints adapter-ready, democratizing scalable inference for all models.
title PHLoRA: data-free Post-hoc Low-Rank Adapter extraction from full-rank checkpoint
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.10971