EmbedLLM: Learning Compact Representations of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhuang, Richard, Wu, Tianhao, Wen, Zhaojin, Li, Andrew, Jiao, Jiantao, Ramchandran, Kannan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929547350376448
author Zhuang, Richard
Wu, Tianhao
Wen, Zhaojin
Li, Andrew
Jiao, Jiantao
Ramchandran, Kannan
author_facet Zhuang, Richard
Wu, Tianhao
Wen, Zhaojin
Li, Andrew
Jiao, Jiantao
Ramchandran, Kannan
contents With hundreds of thousands of language models available on Huggingface today, efficiently evaluating and utilizing these models across various downstream, tasks has become increasingly critical. Many existing methods repeatedly learn task-specific representations of Large Language Models (LLMs), which leads to inefficiencies in both time and computational resources. To address this, we propose EmbedLLM, a framework designed to learn compact vector representations, of LLMs that facilitate downstream applications involving many models, such as model routing. We introduce an encoder-decoder approach for learning such embeddings, along with a systematic framework to evaluate their effectiveness. Empirical results show that EmbedLLM outperforms prior methods in model routing both in accuracy and latency. Additionally, we demonstrate that our method can forecast a model's performance on multiple benchmarks, without incurring additional inference cost. Extensive probing experiments validate that the learned embeddings capture key model characteristics, e.g. whether the model is specialized for coding tasks, even without being explicitly trained on them. We open source our dataset, code and embedder to facilitate further research and application.
format Preprint
id arxiv_https___arxiv_org_abs_2410_02223
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EmbedLLM: Learning Compact Representations of Large Language Models
Zhuang, Richard
Wu, Tianhao
Wen, Zhaojin
Li, Andrew
Jiao, Jiantao
Ramchandran, Kannan
Computation and Language
Artificial Intelligence
Machine Learning
With hundreds of thousands of language models available on Huggingface today, efficiently evaluating and utilizing these models across various downstream, tasks has become increasingly critical. Many existing methods repeatedly learn task-specific representations of Large Language Models (LLMs), which leads to inefficiencies in both time and computational resources. To address this, we propose EmbedLLM, a framework designed to learn compact vector representations, of LLMs that facilitate downstream applications involving many models, such as model routing. We introduce an encoder-decoder approach for learning such embeddings, along with a systematic framework to evaluate their effectiveness. Empirical results show that EmbedLLM outperforms prior methods in model routing both in accuracy and latency. Additionally, we demonstrate that our method can forecast a model's performance on multiple benchmarks, without incurring additional inference cost. Extensive probing experiments validate that the learned embeddings capture key model characteristics, e.g. whether the model is specialized for coding tasks, even without being explicitly trained on them. We open source our dataset, code and embedder to facilitate further research and application.
title EmbedLLM: Learning Compact Representations of Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.02223