Compact Recurrent Transformer with Persistent Memory
Fuente:
arXiv
Salvato in:
| Autori principali: | Mucllari, Edison, Daniels, Zachary, Zhang, David, Ye, Qiang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Noise-Tolerant Coreset-Based Class Incremental Continual Learning
di: Mucllari, Edison, et al.
Pubblicazione: (2025)
di: Mucllari, Edison, et al.
Pubblicazione: (2025)
Recurrent Action Transformer with Memory
di: Cherepanov, Egor, et al.
Pubblicazione: (2023)
di: Cherepanov, Egor, et al.
Pubblicazione: (2023)
Associative Recurrent Memory Transformer
di: Rodkin, Ivan, et al.
Pubblicazione: (2024)
di: Rodkin, Ivan, et al.
Pubblicazione: (2024)
Renaissance of RNNs in Streaming Clinical Time Series: Compact Recurrence Remains Competitive with Transformers
di: Tong, Ran, et al.
Pubblicazione: (2025)
di: Tong, Ran, et al.
Pubblicazione: (2025)
Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
di: Sivtsov, Danil, et al.
Pubblicazione: (2025)
di: Sivtsov, Danil, et al.
Pubblicazione: (2025)
Latent Context Compilation: Distilling Long Context into Compact Portable Memory
di: Li, Zeju, et al.
Pubblicazione: (2026)
di: Li, Zeju, et al.
Pubblicazione: (2026)
Compact Memory for Continual Logistic Regression
di: Jung, Yohan, et al.
Pubblicazione: (2025)
di: Jung, Yohan, et al.
Pubblicazione: (2025)
Towards Compressive and Scalable Recurrent Memory
di: Song, Yunchong, et al.
Pubblicazione: (2026)
di: Song, Yunchong, et al.
Pubblicazione: (2026)
RP-CATE: Recurrent Perceptron-based Channel Attention Transformer Encoder for Industrial Hybrid Modeling
di: Yang, Haoran, et al.
Pubblicazione: (2025)
di: Yang, Haoran, et al.
Pubblicazione: (2025)
Memory Capacity of Nonlinear Recurrent Networks: Is it Informative?
di: Ballarin, Giovanni, et al.
Pubblicazione: (2025)
di: Ballarin, Giovanni, et al.
Pubblicazione: (2025)
Recurrent Memory for Online Interdomain Gaussian Processes
di: Chen, Wenlong, et al.
Pubblicazione: (2025)
di: Chen, Wenlong, et al.
Pubblicazione: (2025)
Memori: A Persistent Memory Layer for Efficient, Context-Aware LLM Agents
di: Borro, Luiz C., et al.
Pubblicazione: (2026)
di: Borro, Luiz C., et al.
Pubblicazione: (2026)
GPU Memory Requirement Prediction for Deep Learning Task Based on Bidirectional Gated Recurrent Unit Optimization Transformer
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
Convolutional Persistence Transforms
di: Solomon, Elchanan, et al.
Pubblicazione: (2022)
di: Solomon, Elchanan, et al.
Pubblicazione: (2022)
Transformer Key-Value Memories Are Nearly as Interpretable as Sparse Autoencoders
di: Ye, Mengyu, et al.
Pubblicazione: (2025)
di: Ye, Mengyu, et al.
Pubblicazione: (2025)
Understanding Dynamic Compute Allocation in Recurrent Transformers
di: Moosa, Ibraheem Muhammad, et al.
Pubblicazione: (2026)
di: Moosa, Ibraheem Muhammad, et al.
Pubblicazione: (2026)
Learning in Compact Spaces with Approximately Normalized Transformer
di: Franke, Jörg K. H., et al.
Pubblicazione: (2025)
di: Franke, Jörg K. H., et al.
Pubblicazione: (2025)
Separations in the Representational Capabilities of Transformers and Recurrent Architectures
di: Bhattamishra, Satwik, et al.
Pubblicazione: (2024)
di: Bhattamishra, Satwik, et al.
Pubblicazione: (2024)
T-SHRED: Symbolic Regression for Regularization and Model Discovery with Transformer Shallow Recurrent Decoders
di: Yermakov, Alexey, et al.
Pubblicazione: (2025)
di: Yermakov, Alexey, et al.
Pubblicazione: (2025)
Multiset Transformer: Advancing Representation Learning in Persistence Diagrams
di: Wang, Minghua, et al.
Pubblicazione: (2024)
di: Wang, Minghua, et al.
Pubblicazione: (2024)
Transformation of audio embeddings into interpretable, concept-based representations
di: Zhang, Alice, et al.
Pubblicazione: (2025)
di: Zhang, Alice, et al.
Pubblicazione: (2025)
Trained Persistent Memory for Frozen Decoder-Only LLMs
di: Jeong, Hong
Pubblicazione: (2026)
di: Jeong, Hong
Pubblicazione: (2026)
Two-Scale Latent Dynamics for Recurrent-Depth Transformers
di: Pappone, Francesco, et al.
Pubblicazione: (2025)
di: Pappone, Francesco, et al.
Pubblicazione: (2025)
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
di: Oncescu, Costin-Andrei, et al.
Pubblicazione: (2026)
di: Oncescu, Costin-Andrei, et al.
Pubblicazione: (2026)
On the Contractivity of Stochastic Interpolation Flow
di: Daniels, Mara
Pubblicazione: (2025)
di: Daniels, Mara
Pubblicazione: (2025)
xPerT: Extended Persistence Transformer
di: Kim, Sehun
Pubblicazione: (2024)
di: Kim, Sehun
Pubblicazione: (2024)
READ: Recurrent Adaptation of Large Transformers
di: Nguyen, John, et al.
Pubblicazione: (2023)
di: Nguyen, John, et al.
Pubblicazione: (2023)
Frailty-Aware Transformer for Recurrent Survival Modeling of Driver Retention in Ride-Hailing Platforms
di: Xu, Shuoyan, et al.
Pubblicazione: (2025)
di: Xu, Shuoyan, et al.
Pubblicazione: (2025)
Next-Latent Prediction Transformers Learn Compact World Models
di: Teoh, Jayden, et al.
Pubblicazione: (2025)
di: Teoh, Jayden, et al.
Pubblicazione: (2025)
Decision Trees That Remember: Gradient-Based Learning of Recurrent Decision Trees with Memory
di: Marton, Sascha, et al.
Pubblicazione: (2025)
di: Marton, Sascha, et al.
Pubblicazione: (2025)
PRformer: Pyramidal Recurrent Transformer for Multivariate Time Series Forecasting
di: Yu, Yongbo, et al.
Pubblicazione: (2024)
di: Yu, Yongbo, et al.
Pubblicazione: (2024)
A note on the VC dimension of 1-dimensional GNNs
di: Daniëls, Noah, et al.
Pubblicazione: (2024)
di: Daniëls, Noah, et al.
Pubblicazione: (2024)
Chain-of-Thought and Compressed Looped Transformers: A Memory-Budget Separation
di: Zhang, Haozhou
Pubblicazione: (2026)
di: Zhang, Haozhou
Pubblicazione: (2026)
Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory
di: Goldstein, Daniel, et al.
Pubblicazione: (2026)
di: Goldstein, Daniel, et al.
Pubblicazione: (2026)
Quantitative Bounds for Length Generalization in Transformers
di: Izzo, Zachary, et al.
Pubblicazione: (2025)
di: Izzo, Zachary, et al.
Pubblicazione: (2025)
Memory Limitations of Prompt Tuning in Transformers
di: Meyer, Maxime, et al.
Pubblicazione: (2025)
di: Meyer, Maxime, et al.
Pubblicazione: (2025)
Recurrent Memory-Augmented Transformers with Chunked Attention for Long-Context Language Modeling
di: Kashyap, Ankit
Pubblicazione: (2025)
di: Kashyap, Ankit
Pubblicazione: (2025)
Why Linear Recurrent Memory Works in Partially Observable Reinforcement Learning
di: Zhao, Yike, et al.
Pubblicazione: (2026)
di: Zhao, Yike, et al.
Pubblicazione: (2026)
Modeling Long Sequences in Bladder Cancer Recurrence: A Comparative Evaluation of LSTM,Transformer,and Mamba
di: Zhang, Runquan, et al.
Pubblicazione: (2024)
di: Zhang, Runquan, et al.
Pubblicazione: (2024)
TRecViT: A Recurrent Video Transformer
di: Pătrăucean, Viorica, et al.
Pubblicazione: (2024)
di: Pătrăucean, Viorica, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Noise-Tolerant Coreset-Based Class Incremental Continual Learning
di: Mucllari, Edison, et al.
Pubblicazione: (2025) -
Recurrent Action Transformer with Memory
di: Cherepanov, Egor, et al.
Pubblicazione: (2023) -
Associative Recurrent Memory Transformer
di: Rodkin, Ivan, et al.
Pubblicazione: (2024) -
Renaissance of RNNs in Streaming Clinical Time Series: Compact Recurrence Remains Competitive with Transformers
di: Tong, Ran, et al.
Pubblicazione: (2025) -
Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
di: Sivtsov, Danil, et al.
Pubblicazione: (2025)