ConsistentEE: A Consistent and Hardness-Guided Early Exiting Method for Accelerating Language Models Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Ziqian, Hong, Yihuai, Dai, Hongliang, Zhuang, Huiping, Chen, Cen
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914743036411904
author Zeng, Ziqian
Hong, Yihuai
Dai, Hongliang
Zhuang, Huiping
Chen, Cen
author_facet Zeng, Ziqian
Hong, Yihuai
Dai, Hongliang
Zhuang, Huiping
Chen, Cen
contents Early Exiting is one of the most popular methods to achieve efficient inference. Current early exiting methods adopt the (weighted) sum of the cross entropy loss of all internal classifiers during training, imposing all these classifiers to predict all instances correctly. However, during inference, as long as one internal classifier predicts an instance correctly, it can accelerate without losing accuracy. Thus, there is a notable gap between training and inference. We propose ConsistentEE, an early exiting method that is consistent in training and inference. ConsistentEE formulates the early exiting process as a reinforcement learning problem. A policy network is added to decide whether an instance should exit or continue. The training objective of ConsistentEE only require each instance to be predicted correctly by one internal classifier. Additionally, we introduce the concept Memorize Layer to measure the hardness of an instance. We incorporate memorized layer into reward function design, which allows "easy" instances to focus more on acceleration while "hard" instances to focus more on accuracy. Experimental results show that our method outperforms other baselines on various natural language understanding and generation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2312_11882
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle ConsistentEE: A Consistent and Hardness-Guided Early Exiting Method for Accelerating Language Models Inference
Zeng, Ziqian
Hong, Yihuai
Dai, Hongliang
Zhuang, Huiping
Chen, Cen
Computation and Language
Artificial Intelligence
Machine Learning
Early Exiting is one of the most popular methods to achieve efficient inference. Current early exiting methods adopt the (weighted) sum of the cross entropy loss of all internal classifiers during training, imposing all these classifiers to predict all instances correctly. However, during inference, as long as one internal classifier predicts an instance correctly, it can accelerate without losing accuracy. Thus, there is a notable gap between training and inference. We propose ConsistentEE, an early exiting method that is consistent in training and inference. ConsistentEE formulates the early exiting process as a reinforcement learning problem. A policy network is added to decide whether an instance should exit or continue. The training objective of ConsistentEE only require each instance to be predicted correctly by one internal classifier. Additionally, we introduce the concept Memorize Layer to measure the hardness of an instance. We incorporate memorized layer into reward function design, which allows "easy" instances to focus more on acceleration while "hard" instances to focus more on accuracy. Experimental results show that our method outperforms other baselines on various natural language understanding and generation tasks.
title ConsistentEE: A Consistent and Hardness-Guided Early Exiting Method for Accelerating Language Models Inference
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2312.11882