CharED: Character-wise Ensemble Decoding for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gu, Kevin, Tuecke, Eva, Katz, Dmitriy, Horesh, Raya, Alvarez-Melis, David, Yurochkin, Mikhail
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914871913742336
author Gu, Kevin
Tuecke, Eva
Katz, Dmitriy
Horesh, Raya
Alvarez-Melis, David
Yurochkin, Mikhail
author_facet Gu, Kevin
Tuecke, Eva
Katz, Dmitriy
Horesh, Raya
Alvarez-Melis, David
Yurochkin, Mikhail
contents Large language models (LLMs) have shown remarkable potential for problem solving, with open source models achieving increasingly impressive performance on benchmarks measuring areas from logical reasoning to mathematical ability. Ensembling models can further improve capabilities across a variety of domains. However, conventional methods of combining models at inference time such as shallow fusion necessitate a shared vocabulary and tokenization, and alternatives like fine-tuning for domain-specific performance are both time consuming and computationally expensive. We therefore present an inference-time ensembling algorithm aimed at "averaging" outputs from multiple LLMs and illustrate its improved performance across multiple domains compared to its constituent models alone. Character-wise ensemble decoding, CharED, finds the marginal distribution of each character for an individual model and performs a weighted average to generate an output, character by character. In coding, math, and toxicity benchmarks, we find our proposed model able to combine complimentary strengths of multiple LLMs, regardless of vocabulary, tokenization, or model size.
format Preprint
id arxiv_https___arxiv_org_abs_2407_11009
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CharED: Character-wise Ensemble Decoding for Large Language Models
Gu, Kevin
Tuecke, Eva
Katz, Dmitriy
Horesh, Raya
Alvarez-Melis, David
Yurochkin, Mikhail
Computation and Language
Machine Learning
Large language models (LLMs) have shown remarkable potential for problem solving, with open source models achieving increasingly impressive performance on benchmarks measuring areas from logical reasoning to mathematical ability. Ensembling models can further improve capabilities across a variety of domains. However, conventional methods of combining models at inference time such as shallow fusion necessitate a shared vocabulary and tokenization, and alternatives like fine-tuning for domain-specific performance are both time consuming and computationally expensive. We therefore present an inference-time ensembling algorithm aimed at "averaging" outputs from multiple LLMs and illustrate its improved performance across multiple domains compared to its constituent models alone. Character-wise ensemble decoding, CharED, finds the marginal distribution of each character for an individual model and performs a weighted average to generate an output, character by character. In coding, math, and toxicity benchmarks, we find our proposed model able to combine complimentary strengths of multiple LLMs, regardless of vocabulary, tokenization, or model size.
title CharED: Character-wise Ensemble Decoding for Large Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2407.11009