You Only Scan Once: Efficient Multi-dimension Sequential Modeling with LightNet

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qin, Zhen, Mao, Yuxin, Shen, Xuyang, Li, Dong, Zhang, Jing, Dai, Yuchao, Zhong, Yiran
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909214067130368
author Qin, Zhen
Mao, Yuxin
Shen, Xuyang
Li, Dong
Zhang, Jing
Dai, Yuchao
Zhong, Yiran
author_facet Qin, Zhen
Mao, Yuxin
Shen, Xuyang
Li, Dong
Zhang, Jing
Dai, Yuchao
Zhong, Yiran
contents Linear attention mechanisms have gained prominence in causal language models due to their linear computational complexity and enhanced speed. However, the inherent decay mechanism in linear attention presents challenges when applied to multi-dimensional sequence modeling tasks, such as image processing and multi-modal learning. In these scenarios, the utilization of sequential scanning to establish a global receptive field necessitates multiple scans for multi-dimensional data, thereby leading to inefficiencies. This paper identifies the inefficiency caused by a multiplicative linear recurrence and proposes an efficient alternative additive linear recurrence to avoid the issue, as it can handle multi-dimensional data within a single scan. We further develop an efficient multi-dimensional sequential modeling framework called LightNet based on the new recurrence. Moreover, we present two new multi-dimensional linear relative positional encoding methods, MD-TPE and MD-LRPE to enhance the model's ability to discern positional information in multi-dimensional scenarios. Our empirical evaluations across various tasks, including image classification, image generation, bidirectional language modeling, and autoregressive language modeling, demonstrate the efficacy of LightNet, showcasing its potential as a versatile and efficient solution for multi-dimensional sequential modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2405_21022
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle You Only Scan Once: Efficient Multi-dimension Sequential Modeling with LightNet
Qin, Zhen
Mao, Yuxin
Shen, Xuyang
Li, Dong
Zhang, Jing
Dai, Yuchao
Zhong, Yiran
Computation and Language
Computer Vision and Pattern Recognition
Linear attention mechanisms have gained prominence in causal language models due to their linear computational complexity and enhanced speed. However, the inherent decay mechanism in linear attention presents challenges when applied to multi-dimensional sequence modeling tasks, such as image processing and multi-modal learning. In these scenarios, the utilization of sequential scanning to establish a global receptive field necessitates multiple scans for multi-dimensional data, thereby leading to inefficiencies. This paper identifies the inefficiency caused by a multiplicative linear recurrence and proposes an efficient alternative additive linear recurrence to avoid the issue, as it can handle multi-dimensional data within a single scan. We further develop an efficient multi-dimensional sequential modeling framework called LightNet based on the new recurrence. Moreover, we present two new multi-dimensional linear relative positional encoding methods, MD-TPE and MD-LRPE to enhance the model's ability to discern positional information in multi-dimensional scenarios. Our empirical evaluations across various tasks, including image classification, image generation, bidirectional language modeling, and autoregressive language modeling, demonstrate the efficacy of LightNet, showcasing its potential as a versatile and efficient solution for multi-dimensional sequential modeling.
title You Only Scan Once: Efficient Multi-dimension Sequential Modeling with LightNet
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.21022