Herd: Using multiple, smaller LLMs to match the performances of proprietary, large LLMs via an intelligent composer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hari, Surya Narayanan, Liu, Rex, Thomson, Matt
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914954023534592
author Hari, Surya Narayanan
Liu, Rex
Thomson, Matt
author_facet Hari, Surya Narayanan
Liu, Rex
Thomson, Matt
contents Currently, over a thousand LLMs exist that are multi-purpose and are capable of performing real world tasks, including Q&A, text summarization, content generation, etc. However, accessibility, scale and reliability of free models prevents them from being widely deployed in everyday use cases. To address the first two issues of access and scale, organisations such as HuggingFace have created model repositories where users have uploaded model weights and quantized versions of models trained using different paradigms, as well as model cards describing their training process. While some models report performance on commonly used benchmarks, not all do, and interpreting the real world impact of trading off performance on a benchmark for model deployment cost, is unclear. Here, we show that a herd of open source models can match or exceed the performance of proprietary models via an intelligent router. We show that a Herd of open source models is able to match the accuracy of ChatGPT, despite being composed of models that are effectively 2.5x smaller. We show that in cases where GPT is not able to answer the query, Herd is able to identify a model that can, at least 40% of the time.
format Preprint
id arxiv_https___arxiv_org_abs_2310_19902
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Herd: Using multiple, smaller LLMs to match the performances of proprietary, large LLMs via an intelligent composer
Hari, Surya Narayanan
Liu, Rex
Thomson, Matt
Artificial Intelligence
Currently, over a thousand LLMs exist that are multi-purpose and are capable of performing real world tasks, including Q&A, text summarization, content generation, etc. However, accessibility, scale and reliability of free models prevents them from being widely deployed in everyday use cases. To address the first two issues of access and scale, organisations such as HuggingFace have created model repositories where users have uploaded model weights and quantized versions of models trained using different paradigms, as well as model cards describing their training process. While some models report performance on commonly used benchmarks, not all do, and interpreting the real world impact of trading off performance on a benchmark for model deployment cost, is unclear. Here, we show that a herd of open source models can match or exceed the performance of proprietary models via an intelligent router. We show that a Herd of open source models is able to match the accuracy of ChatGPT, despite being composed of models that are effectively 2.5x smaller. We show that in cases where GPT is not able to answer the query, Herd is able to identify a model that can, at least 40% of the time.
title Herd: Using multiple, smaller LLMs to match the performances of proprietary, large LLMs via an intelligent composer
topic Artificial Intelligence
url https://arxiv.org/abs/2310.19902