Vision-LSTM: xLSTM as Generic Vision Backbone
Fuente:
arXiv
Saved in:
| Main Authors: | Alkin, Benedikt, Beck, Maximilian, Pöppel, Korbinian, Hochreiter, Sepp, Brandstetter, Johannes |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tiled Flash Linear Attention: More Efficient Linear RNN and xLSTM Kernels
by: Beck, Maximilian, et al.
Published: (2025)
by: Beck, Maximilian, et al.
Published: (2025)
xLSTM: Extended Long Short-Term Memory
by: Beck, Maximilian, et al.
Published: (2024)
by: Beck, Maximilian, et al.
Published: (2024)
MIM-Refiner: A Contrastive Learning Boost from Intermediate Pre-Trained Representations
by: Alkin, Benedikt, et al.
Published: (2024)
by: Alkin, Benedikt, et al.
Published: (2024)
A Large Recurrent Action Model: xLSTM enables Fast Inference for Robotics Tasks
by: Schmied, Thomas, et al.
Published: (2024)
by: Schmied, Thomas, et al.
Published: (2024)
xLSTM 7B: A Recurrent LLM for Fast and Efficient Inference
by: Beck, Maximilian, et al.
Published: (2025)
by: Beck, Maximilian, et al.
Published: (2025)
Seg-LSTM: Performance of xLSTM for Semantic Segmentation of Remotely Sensed Images
by: Zhu, Qinfeng, et al.
Published: (2024)
by: Zhu, Qinfeng, et al.
Published: (2024)
FlashRNN: I/O-Aware Optimization of Traditional RNNs on modern hardware
by: Pöppel, Korbinian, et al.
Published: (2024)
by: Pöppel, Korbinian, et al.
Published: (2024)
Bio-xLSTM: Generative modeling, representation and in-context learning of biological and chemical sequences
by: Schmidinger, Niklas, et al.
Published: (2024)
by: Schmidinger, Niklas, et al.
Published: (2024)
xLSTM Scaling Laws: Competitive Performance with Linear Time-Complexity
by: Beck, Maximilian, et al.
Published: (2025)
by: Beck, Maximilian, et al.
Published: (2025)
ViR: Towards Efficient Vision Retention Backbones
by: Hatamizadeh, Ali, et al.
Published: (2023)
by: Hatamizadeh, Ali, et al.
Published: (2023)
Effective Distillation to Hybrid xLSTM Architectures
by: Hauzenberger, Lukas, et al.
Published: (2026)
by: Hauzenberger, Lukas, et al.
Published: (2026)
xLSTM-ECG: Multi-label ECG Classification via Feature Fusion with xLSTM
by: Kang, Lei, et al.
Published: (2025)
by: Kang, Lei, et al.
Published: (2025)
Geometry-Informed Neural Networks
by: Berzins, Arturs, et al.
Published: (2024)
by: Berzins, Arturs, et al.
Published: (2024)
Vision Backbone Efficient Selection for Image Classification in Low-Data Regimes
by: Guerin, Joris, et al.
Published: (2024)
by: Guerin, Joris, et al.
Published: (2024)
SAFE-KD: Risk-Controlled Early-Exit Distillation for Vision Backbones
by: Khazem, Salim
Published: (2026)
by: Khazem, Salim
Published: (2026)
MAL: Cluster-Masked and Multi-Task Pretraining for Enhanced xLSTM Vision Performance
by: Huang, Wenjun, et al.
Published: (2024)
by: Huang, Wenjun, et al.
Published: (2024)
Linear Alignment of Vision-language Models for Image Captioning
by: Paischer, Fabian, et al.
Published: (2023)
by: Paischer, Fabian, et al.
Published: (2023)
pLSTM: parallelizable Linear Source Transition Mark networks
by: Pöppel, Korbinian, et al.
Published: (2025)
by: Pöppel, Korbinian, et al.
Published: (2025)
ASTM :Autonomous Smart Traffic Management System Using Artificial Intelligence CNN and LSTM
by: Goenawan, Christofel Rio
Published: (2024)
by: Goenawan, Christofel Rio
Published: (2024)
xLSTM-UNet can be an Effective 2D & 3D Medical Image Segmentation Backbone with Vision-LSTM (ViL) better than its Mamba Counterpart
by: Chen, Tianrun, et al.
Published: (2024)
by: Chen, Tianrun, et al.
Published: (2024)
Distil-xLSTM: Learning Attention Mechanisms through Recurrent Structures
by: Thiombiano, Abdoul Majid O., et al.
Published: (2025)
by: Thiombiano, Abdoul Majid O., et al.
Published: (2025)
Channel Vision Transformers: An Image Is Worth 1 x 16 x 16 Words
by: Bao, Yujia, et al.
Published: (2023)
by: Bao, Yujia, et al.
Published: (2023)
Comparative Analysis of Liquid Neural Networks and LSTM for Sequential Pattern Recognition: Robustness, Efficiency, and Clinical Utility
by: Thu, Ye Kyaw, et al.
Published: (2026)
by: Thu, Ye Kyaw, et al.
Published: (2026)
xLSTM-FER: Enhancing Student Expression Recognition with Extended Vision Long Short-Term Memory Network
by: Huang, Qionghao, et al.
Published: (2024)
by: Huang, Qionghao, et al.
Published: (2024)
Predicting Blastocyst Formation in IVF: Integrating DINOv2 and Attention-Based LSTM on Time-Lapse Embryo Images
by: Varzaneh, Zahra Asghari, et al.
Published: (2026)
by: Varzaneh, Zahra Asghari, et al.
Published: (2026)
Evaluating BiLSTM and CNN+GRU Approaches for Human Activity Recognition Using WiFi CSI Data
by: Wakili, Almustapha A., et al.
Published: (2025)
by: Wakili, Almustapha A., et al.
Published: (2025)
Improving Urban Flood Prediction using LSTM-DeepLabv3+ and Bayesian Optimization with Spatiotemporal feature fusion
by: Situ, Zuxiang, et al.
Published: (2023)
by: Situ, Zuxiang, et al.
Published: (2023)
Shape Generation via Weight Space Learning
by: Plattner, Maximilian, et al.
Published: (2025)
by: Plattner, Maximilian, et al.
Published: (2025)
xLSTMTime : Long-term Time Series Forecasting With xLSTM
by: Alharthi, Musleh, et al.
Published: (2024)
by: Alharthi, Musleh, et al.
Published: (2024)
Enhancing Spatiotemporal Networks with xLSTM: A Scalar LSTM Approach for Cellular Traffic Forecasting
by: Ali, Khalid, et al.
Published: (2025)
by: Ali, Khalid, et al.
Published: (2025)
Are Vision xLSTM Embedded UNet More Reliable in Medical 3D Image Segmentation?
by: Dutta, Pallabi, et al.
Published: (2024)
by: Dutta, Pallabi, et al.
Published: (2024)
When Better Eyes Lead to Blindness: A Diagnostic Study of the Information Bottleneck in CNN-LSTM Image Captioning Models
by: Gupta, Hitesh Kumar
Published: (2025)
by: Gupta, Hitesh Kumar
Published: (2025)
VariViT: A Vision Transformer for Variable Image Sizes
by: Varma, Aswathi, et al.
Published: (2026)
by: Varma, Aswathi, et al.
Published: (2026)
Generalized Category Discovery under Domain Shifts: From Vision to Vision-Language Models
by: Wang, Hongjun, et al.
Published: (2026)
by: Wang, Hongjun, et al.
Published: (2026)
EfficientTrain++: Generalized Curriculum Learning for Efficient Visual Backbone Training
by: Wang, Yulin, et al.
Published: (2024)
by: Wang, Yulin, et al.
Published: (2024)
An Overview of Prototype Formulations for Interpretable Deep Learning
by: Li, Maximilian Xiling, et al.
Published: (2024)
by: Li, Maximilian Xiling, et al.
Published: (2024)
Diverse Topology Optimization using Modulated Neural Fields
by: Radler, Andreas, et al.
Published: (2025)
by: Radler, Andreas, et al.
Published: (2025)
Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf Interactions
by: Baniecki, Hubert, et al.
Published: (2025)
by: Baniecki, Hubert, et al.
Published: (2025)
xLSTMAD: A Powerful xLSTM-based Method for Anomaly Detection
by: Faber, Kamil, et al.
Published: (2025)
by: Faber, Kamil, et al.
Published: (2025)
$MV_{Hybrid}$: Improving Spatial Transcriptomics Prediction with Hybrid State Space-Vision Transformer Backbone in Pathology Vision Foundation Models
by: Cho, Won June, et al.
Published: (2025)
by: Cho, Won June, et al.
Published: (2025)
Similar Items
-
Tiled Flash Linear Attention: More Efficient Linear RNN and xLSTM Kernels
by: Beck, Maximilian, et al.
Published: (2025) -
xLSTM: Extended Long Short-Term Memory
by: Beck, Maximilian, et al.
Published: (2024) -
MIM-Refiner: A Contrastive Learning Boost from Intermediate Pre-Trained Representations
by: Alkin, Benedikt, et al.
Published: (2024) -
A Large Recurrent Action Model: xLSTM enables Fast Inference for Robotics Tasks
by: Schmied, Thomas, et al.
Published: (2024) -
xLSTM 7B: A Recurrent LLM for Fast and Efficient Inference
by: Beck, Maximilian, et al.
Published: (2025)