Iterative Forward Tuning Boosts In-Context Learning in Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Jiaxi, Hui, Binyuan, Yang, Min, Wang, Bailin, Li, Bowen, Li, Binhua, Huang, Fei, Li, Yongbin
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911901110239232
author Yang, Jiaxi
Hui, Binyuan
Yang, Min
Wang, Bailin
Li, Bowen
Li, Binhua
Huang, Fei
Li, Yongbin
author_facet Yang, Jiaxi
Hui, Binyuan
Yang, Min
Wang, Bailin
Li, Bowen
Li, Binhua
Huang, Fei
Li, Yongbin
contents Despite the advancements in in-context learning (ICL) for large language models (LLMs), current research centers on specific prompt engineering, such as demonstration selection, with the expectation that a single iteration of demonstrations processing can generalize effectively to a given test sample. However, this perspective overlooks the potential benefits derived from multiple iterations involving demonstrations, a practice aligning more closely with the iterative decision-making process exhibited by humans, who often learn through analogy. In this study, we introduce a novel two-stage framework to boost ICL in LLMs. Specifically, our framework delineates the ICL process into two distinct stages: Deep-Thinking and test stages. The Deep-Thinking stage incorporates a unique attention mechanism, i.e., iterative enhanced attention, which enables multiple rounds of information accumulation. This mechanism operates by manipulating the Key-Value matrices without training, fostering enhanced understanding capabilities in LLMs by thinking demonstrations multiple times. We evaluated Deep-Thinking across a range of benchmarks and LLMs, showing its superior performance over vanilla ICL methods and its effectiveness in challenging tasks where demonstration selection is infeasible.
format Preprint
id arxiv_https___arxiv_org_abs_2305_13016
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Iterative Forward Tuning Boosts In-Context Learning in Language Models
Yang, Jiaxi
Hui, Binyuan
Yang, Min
Wang, Bailin
Li, Bowen
Li, Binhua
Huang, Fei
Li, Yongbin
Computation and Language
Despite the advancements in in-context learning (ICL) for large language models (LLMs), current research centers on specific prompt engineering, such as demonstration selection, with the expectation that a single iteration of demonstrations processing can generalize effectively to a given test sample. However, this perspective overlooks the potential benefits derived from multiple iterations involving demonstrations, a practice aligning more closely with the iterative decision-making process exhibited by humans, who often learn through analogy. In this study, we introduce a novel two-stage framework to boost ICL in LLMs. Specifically, our framework delineates the ICL process into two distinct stages: Deep-Thinking and test stages. The Deep-Thinking stage incorporates a unique attention mechanism, i.e., iterative enhanced attention, which enables multiple rounds of information accumulation. This mechanism operates by manipulating the Key-Value matrices without training, fostering enhanced understanding capabilities in LLMs by thinking demonstrations multiple times. We evaluated Deep-Thinking across a range of benchmarks and LLMs, showing its superior performance over vanilla ICL methods and its effectiveness in challenging tasks where demonstration selection is infeasible.
title Iterative Forward Tuning Boosts In-Context Learning in Language Models
topic Computation and Language
url https://arxiv.org/abs/2305.13016