Vision-LSTM: xLSTM as Generic Vision Backbone

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Alkin, Benedikt, Beck, Maximilian, Pöppel, Korbinian, Hochreiter, Sepp, Brandstetter, Johannes
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912239509831680
author Alkin, Benedikt
Beck, Maximilian
Pöppel, Korbinian
Hochreiter, Sepp
Brandstetter, Johannes
author_facet Alkin, Benedikt
Beck, Maximilian
Pöppel, Korbinian
Hochreiter, Sepp
Brandstetter, Johannes
contents Transformers are widely used as generic backbones in computer vision, despite initially introduced for natural language processing. Recently, the Long Short-Term Memory (LSTM) has been extended to a scalable and performant architecture - the xLSTM - which overcomes long-standing LSTM limitations via exponential gating and parallelizable matrix memory structure. In this report, we introduce Vision-LSTM (ViL), an adaption of the xLSTM building blocks to computer vision. ViL comprises a stack of xLSTM blocks where odd blocks process the sequence of patch tokens from top to bottom while even blocks go from bottom to top. Experiments show that ViL holds promise to be further deployed as new generic backbone for computer vision architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2406_04303
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Vision-LSTM: xLSTM as Generic Vision Backbone
Alkin, Benedikt
Beck, Maximilian
Pöppel, Korbinian
Hochreiter, Sepp
Brandstetter, Johannes
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Transformers are widely used as generic backbones in computer vision, despite initially introduced for natural language processing. Recently, the Long Short-Term Memory (LSTM) has been extended to a scalable and performant architecture - the xLSTM - which overcomes long-standing LSTM limitations via exponential gating and parallelizable matrix memory structure. In this report, we introduce Vision-LSTM (ViL), an adaption of the xLSTM building blocks to computer vision. ViL comprises a stack of xLSTM blocks where odd blocks process the sequence of patch tokens from top to bottom while even blocks go from bottom to top. Experiments show that ViL holds promise to be further deployed as new generic backbone for computer vision architectures.
title Vision-LSTM: xLSTM as Generic Vision Backbone
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2406.04303