Revealing the Inherent Instructability of Pre-Trained Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: An, Seokhyun, Kim, Minji, Kim, Hyounghun
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911151714992128
author An, Seokhyun
Kim, Minji
Kim, Hyounghun
author_facet An, Seokhyun
Kim, Minji
Kim, Hyounghun
contents Instruction tuning -- supervised fine-tuning using instruction-response pairs -- is a key step in making pre-trained large language models (LLMs) instructable. Meanwhile, LLMs perform multitask learning during their pre-training, acquiring extensive knowledge and capabilities. We hypothesize that the pre-training stage can enable them to develop the ability to comprehend and address instructions. To verify this, we propose Response Tuning (RT), which removes the instruction and its corresponding mapping to the response from instruction tuning. Instead, it focuses solely on establishing a response distribution. Our experiments demonstrate that RT models, trained only on responses, can effectively respond to a wide range of instructions akin to their instruction-tuned counterparts. In addition, we observe that the models can recognize and reject unsafe queries after learning a safety policy only from the response data. Furthermore, we find that these observations extend to an in-context learning setting. These findings support our hypothesis, highlighting the extensive inherent capabilities of pre-trained LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2410_02465
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Revealing the Inherent Instructability of Pre-Trained Language Models
An, Seokhyun
Kim, Minji
Kim, Hyounghun
Computation and Language
Artificial Intelligence
Instruction tuning -- supervised fine-tuning using instruction-response pairs -- is a key step in making pre-trained large language models (LLMs) instructable. Meanwhile, LLMs perform multitask learning during their pre-training, acquiring extensive knowledge and capabilities. We hypothesize that the pre-training stage can enable them to develop the ability to comprehend and address instructions. To verify this, we propose Response Tuning (RT), which removes the instruction and its corresponding mapping to the response from instruction tuning. Instead, it focuses solely on establishing a response distribution. Our experiments demonstrate that RT models, trained only on responses, can effectively respond to a wide range of instructions akin to their instruction-tuned counterparts. In addition, we observe that the models can recognize and reject unsafe queries after learning a safety policy only from the response data. Furthermore, we find that these observations extend to an in-context learning setting. These findings support our hypothesis, highlighting the extensive inherent capabilities of pre-trained LLMs.
title Revealing the Inherent Instructability of Pre-Trained Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.02465