FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model Deployment

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zaccone, Riccardo, Laskaridis, Stefanos, Ciccone, Marco, Horváth, Samuel
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918531589734400
author Zaccone, Riccardo
Laskaridis, Stefanos
Ciccone, Marco
Horváth, Samuel
author_facet Zaccone, Riccardo
Laskaridis, Stefanos
Ciccone, Marco
Horváth, Samuel
contents The growing scale of deep neural networks, encompassing large language models (LLMs) and vision transformers (ViTs), has made training from scratch prohibitively expensive and deployment increasingly costly. These models are often used as computational monoliths with fixed cost, hindering adaptive deployment across different cost budgets.We argue that nested components, ordered by importance, can be extracted from pretrained models and selectively activated within the available computational budget. To this end, our proposed FlexRank method leverages low-rank weight decomposition with nested, importance-based consolidation to extract submodels of increasing capabilities. Our approach enables a ``train-once, deploy-everywhere'' paradigm offering a graceful trade-off between cost and performance without training from scratch for each budget - advancing practical deployment of large models.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02680
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model Deployment
Zaccone, Riccardo
Laskaridis, Stefanos
Ciccone, Marco
Horváth, Samuel
Machine Learning
The growing scale of deep neural networks, encompassing large language models (LLMs) and vision transformers (ViTs), has made training from scratch prohibitively expensive and deployment increasingly costly. These models are often used as computational monoliths with fixed cost, hindering adaptive deployment across different cost budgets.We argue that nested components, ordered by importance, can be extracted from pretrained models and selectively activated within the available computational budget. To this end, our proposed FlexRank method leverages low-rank weight decomposition with nested, importance-based consolidation to extract submodels of increasing capabilities. Our approach enables a ``train-once, deploy-everywhere'' paradigm offering a graceful trade-off between cost and performance without training from scratch for each budget - advancing practical deployment of large models.
title FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model Deployment
topic Machine Learning
url https://arxiv.org/abs/2602.02680