AutoMix: Automatically Mixing Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Aggarwal, Pranjal, Madaan, Aman, Anand, Ankit, Potharaju, Srividya Pranavi, Mishra, Swaroop, Zhou, Pei, Gupta, Aditya, Rajagopal, Dheeraj, Kappaganthu, Karthik, Yang, Yiming, Upadhyay, Shyam, Faruqui, Manaal, Mausam
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929681113022464
author Aggarwal, Pranjal
Madaan, Aman
Anand, Ankit
Potharaju, Srividya Pranavi
Mishra, Swaroop
Zhou, Pei
Gupta, Aditya
Rajagopal, Dheeraj
Kappaganthu, Karthik
Yang, Yiming
Upadhyay, Shyam
Faruqui, Manaal
Mausam
author_facet Aggarwal, Pranjal
Madaan, Aman
Anand, Ankit
Potharaju, Srividya Pranavi
Mishra, Swaroop
Zhou, Pei
Gupta, Aditya
Rajagopal, Dheeraj
Kappaganthu, Karthik
Yang, Yiming
Upadhyay, Shyam
Faruqui, Manaal
Mausam
contents Large language models (LLMs) are now available from cloud API providers in various sizes and configurations. While this diversity offers a broad spectrum of choices, effectively leveraging the options to optimize computational cost and performance remains challenging. In this work, we present Automix, an approach that strategically routes queries to larger LMs, based on the approximate correctness of outputs from a smaller LM. Central to Automix are two key technical contributions. First, it has a few-shot self-verification mechanism, which estimates the reliability of its own outputs without requiring extensive training. Second, given that self-verification can be noisy, it employs a POMDP based router that can effectively select an appropriately sized model, based on answer confidence. Experiments across five language models and five challenging datasets show that Automix consistently surpasses strong baselines, reducing computational cost by over 50% for comparable performance.
format Preprint
id arxiv_https___arxiv_org_abs_2310_12963
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle AutoMix: Automatically Mixing Language Models
Aggarwal, Pranjal
Madaan, Aman
Anand, Ankit
Potharaju, Srividya Pranavi
Mishra, Swaroop
Zhou, Pei
Gupta, Aditya
Rajagopal, Dheeraj
Kappaganthu, Karthik
Yang, Yiming
Upadhyay, Shyam
Faruqui, Manaal
Mausam
Computation and Language
Artificial Intelligence
Large language models (LLMs) are now available from cloud API providers in various sizes and configurations. While this diversity offers a broad spectrum of choices, effectively leveraging the options to optimize computational cost and performance remains challenging. In this work, we present Automix, an approach that strategically routes queries to larger LMs, based on the approximate correctness of outputs from a smaller LM. Central to Automix are two key technical contributions. First, it has a few-shot self-verification mechanism, which estimates the reliability of its own outputs without requiring extensive training. Second, given that self-verification can be noisy, it employs a POMDP based router that can effectively select an appropriately sized model, based on answer confidence. Experiments across five language models and five challenging datasets show that Automix consistently surpasses strong baselines, reducing computational cost by over 50% for comparable performance.
title AutoMix: Automatically Mixing Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2310.12963