Cascade-Aware Training of Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Congchao, Augenstein, Sean, Rush, Keith, Jitkrittum, Wittawat, Narasimhan, Harikrishna, Rawat, Ankit Singh, Menon, Aditya Krishna, Go, Alec
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911898299006976
author Wang, Congchao
Augenstein, Sean
Rush, Keith
Jitkrittum, Wittawat
Narasimhan, Harikrishna
Rawat, Ankit Singh
Menon, Aditya Krishna
Go, Alec
author_facet Wang, Congchao
Augenstein, Sean
Rush, Keith
Jitkrittum, Wittawat
Narasimhan, Harikrishna
Rawat, Ankit Singh
Menon, Aditya Krishna
Go, Alec
contents Reducing serving cost and latency is a fundamental concern for the deployment of language models (LMs) in business applications. To address this, cascades of LMs offer an effective solution that conditionally employ smaller models for simpler queries. Cascaded systems are typically built with independently trained models, neglecting the advantages of considering inference-time interactions of the cascaded LMs during training. In this paper, we present cascade-aware training(CAT), an approach to optimizing the overall quality-cost performance tradeoff of a cascade of LMs. We achieve inference-time benefits by training the small LM with awareness of its place in a cascade and downstream capabilities. We demonstrate the value of the proposed method with over 60 LM tasks of the SuperGLUE, WMT22, and FLAN2021 datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2406_00060
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Cascade-Aware Training of Language Models
Wang, Congchao
Augenstein, Sean
Rush, Keith
Jitkrittum, Wittawat
Narasimhan, Harikrishna
Rawat, Ankit Singh
Menon, Aditya Krishna
Go, Alec
Computation and Language
Machine Learning
Reducing serving cost and latency is a fundamental concern for the deployment of language models (LMs) in business applications. To address this, cascades of LMs offer an effective solution that conditionally employ smaller models for simpler queries. Cascaded systems are typically built with independently trained models, neglecting the advantages of considering inference-time interactions of the cascaded LMs during training. In this paper, we present cascade-aware training(CAT), an approach to optimizing the overall quality-cost performance tradeoff of a cascade of LMs. We achieve inference-time benefits by training the small LM with awareness of its place in a cascade and downstream capabilities. We demonstrate the value of the proposed method with over 60 LM tasks of the SuperGLUE, WMT22, and FLAN2021 datasets.
title Cascade-Aware Training of Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2406.00060