Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Son, Guijin, Baek, Sangwon, Nam, Sangdae, Jeong, Ilgyun, Kim, Seungone
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909217975173120
author Son, Guijin
Baek, Sangwon
Nam, Sangdae
Jeong, Ilgyun
Kim, Seungone
author_facet Son, Guijin
Baek, Sangwon
Nam, Sangdae
Jeong, Ilgyun
Kim, Seungone
contents Large language models (LLMs) are typically prompted to follow a single instruction per inference call. In this work, we analyze whether LLMs also hold the capability to handle multiple instructions simultaneously, denoted as Multi-Task Inference. For this purpose, we introduce the MTI Bench(Multi-Task Inference Benchmark), a comprehensive evaluation benchmark encompassing 5,000 instances across 25 tasks. Each task in the MTI Bench involves 2 to 3 sub-tasks. As expected, we first demonstrate that Multi-Task Inference reduces the total inference time by 1.46 times in average since it does not require multiple inference calls. Interestingly, contrary to the expectation that LLMs would perform better when tasks are divided, we find that state-of-the-art LLMs, such as Llama-2-Chat-70B and GPT-4, show up to 7.3% and 12.4% improved performance with Multi-Task Inference compared to Single-Task Inference on the MTI Bench. We release the MTI Bench dataset and our code at this link https://github.com/guijinSON/MTI-Bench.
format Preprint
id arxiv_https___arxiv_org_abs_2402_11597
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?
Son, Guijin
Baek, Sangwon
Nam, Sangdae
Jeong, Ilgyun
Kim, Seungone
Computation and Language
Large language models (LLMs) are typically prompted to follow a single instruction per inference call. In this work, we analyze whether LLMs also hold the capability to handle multiple instructions simultaneously, denoted as Multi-Task Inference. For this purpose, we introduce the MTI Bench(Multi-Task Inference Benchmark), a comprehensive evaluation benchmark encompassing 5,000 instances across 25 tasks. Each task in the MTI Bench involves 2 to 3 sub-tasks. As expected, we first demonstrate that Multi-Task Inference reduces the total inference time by 1.46 times in average since it does not require multiple inference calls. Interestingly, contrary to the expectation that LLMs would perform better when tasks are divided, we find that state-of-the-art LLMs, such as Llama-2-Chat-70B and GPT-4, show up to 7.3% and 12.4% improved performance with Multi-Task Inference compared to Single-Task Inference on the MTI Bench. We release the MTI Bench dataset and our code at this link https://github.com/guijinSON/MTI-Bench.
title Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?
topic Computation and Language
url https://arxiv.org/abs/2402.11597