Numerical Investigation of Sequence Modeling Theory using Controllable Memory Functions
Fuente:
arXiv
Salvato in:
| Autori principali: | Jiang, Haotian, Bao, Zeyu, Wang, Shida, Li, Qianxiao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Approximation Rate of the Transformer Architecture for Sequence Modeling
di: Jiang, Haotian, et al.
Pubblicazione: (2023)
di: Jiang, Haotian, et al.
Pubblicazione: (2023)
StableSSM: Alleviating the Curse of Memory in State-space Models through Stable Reparameterization
di: Wang, Shida, et al.
Pubblicazione: (2023)
di: Wang, Shida, et al.
Pubblicazione: (2023)
The Effect of Depth on the Expressivity of Deep Linear State-Space Models
di: Bao, Zeyu, et al.
Pubblicazione: (2025)
di: Bao, Zeyu, et al.
Pubblicazione: (2025)
InfoFlow: A Framework for Multi-Layer Transformer Analysis
di: Yu, Penghao, et al.
Pubblicazione: (2026)
di: Yu, Penghao, et al.
Pubblicazione: (2026)
Inverse Approximation Theory for Nonlinear Recurrent Neural Networks
di: Wang, Shida, et al.
Pubblicazione: (2023)
di: Wang, Shida, et al.
Pubblicazione: (2023)
The Effect of Attention Head Count on Transformer Approximation
di: Yu, Penghao, et al.
Pubblicazione: (2025)
di: Yu, Penghao, et al.
Pubblicazione: (2025)
Accelerating Legacy Numerical Solvers by Non-intrusive Gradient-based Meta-solving
di: Arisaka, Sohei, et al.
Pubblicazione: (2024)
di: Arisaka, Sohei, et al.
Pubblicazione: (2024)
From Generalization Analysis to Optimization Designs for State Space Models
di: Liu, Fusheng, et al.
Pubblicazione: (2024)
di: Liu, Fusheng, et al.
Pubblicazione: (2024)
Autocorrelation Matters: Understanding the Role of Initialization Schemes for State Space Models
di: Liu, Fusheng, et al.
Pubblicazione: (2024)
di: Liu, Fusheng, et al.
Pubblicazione: (2024)
Allocation of Parameters in Transformers
di: Yu, Ruoxi, et al.
Pubblicazione: (2025)
di: Yu, Ruoxi, et al.
Pubblicazione: (2025)
Learning task-specific predictive models for scientific computing
di: Yin, Jianyuan, et al.
Pubblicazione: (2025)
di: Yin, Jianyuan, et al.
Pubblicazione: (2025)
Unifying back-propagation and forward-forward algorithms through model predictive control
di: Ren, Lianhai, et al.
Pubblicazione: (2024)
di: Ren, Lianhai, et al.
Pubblicazione: (2024)
LongSSM: On the Length Extension of State-space Models in Language Modelling
di: Wang, Shida
Pubblicazione: (2024)
di: Wang, Shida
Pubblicazione: (2024)
DynGMA: a robust approach for learning stochastic differential equations from data
di: Zhu, Aiqing, et al.
Pubblicazione: (2024)
di: Zhu, Aiqing, et al.
Pubblicazione: (2024)
Learning Macroscopic Dynamics from Partial Microscopic Observations
di: Chen, Mengyi, et al.
Pubblicazione: (2024)
di: Chen, Mengyi, et al.
Pubblicazione: (2024)
Continuity-Preserving Convolutional Autoencoders for Learning Continuous Latent Dynamical Models from Images
di: Zhu, Aiqing, et al.
Pubblicazione: (2025)
di: Zhu, Aiqing, et al.
Pubblicazione: (2025)
LongVQ: Long Sequence Modeling with Vector Quantization on Structured Memory
di: Liu, Zicheng, et al.
Pubblicazione: (2024)
di: Liu, Zicheng, et al.
Pubblicazione: (2024)
Learning Permutation-invariant Macroscopic Dynamics
di: Han, Zhichao, et al.
Pubblicazione: (2026)
di: Han, Zhichao, et al.
Pubblicazione: (2026)
Machine Unlearning under Retain-Forget Entanglement
di: Cheng, Jingpu, et al.
Pubblicazione: (2026)
di: Cheng, Jingpu, et al.
Pubblicazione: (2026)
On the Curse of Memory in Recurrent Neural Networks: Approximation and Optimization Analysis
di: Li, Zhong, et al.
Pubblicazione: (2020)
di: Li, Zhong, et al.
Pubblicazione: (2020)
On Hypothesis Transfer Learning of Functional Linear Models
di: Lin, Haotian, et al.
Pubblicazione: (2022)
di: Lin, Haotian, et al.
Pubblicazione: (2022)
SMR: State Memory Replay for Long Sequence Modeling
di: Qi, Biqing, et al.
Pubblicazione: (2024)
di: Qi, Biqing, et al.
Pubblicazione: (2024)
Multi-Modal Representation Learning for Molecular Property Prediction: Sequence, Graph, Geometry
di: Wang, Zeyu, et al.
Pubblicazione: (2024)
di: Wang, Zeyu, et al.
Pubblicazione: (2024)
MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration
di: Ren, Lianhai, et al.
Pubblicazione: (2026)
di: Ren, Lianhai, et al.
Pubblicazione: (2026)
Scalable learning of macroscopic stochastic dynamics
di: Chen, Mengyi, et al.
Pubblicazione: (2025)
di: Chen, Mengyi, et al.
Pubblicazione: (2025)
A unified framework for establishing the universal approximation of transformer-type architectures
di: Cheng, Jingpu, et al.
Pubblicazione: (2025)
di: Cheng, Jingpu, et al.
Pubblicazione: (2025)
Deep learning and the rate of approximation by flows
di: Cheng, Jingpu, et al.
Pubblicazione: (2026)
di: Cheng, Jingpu, et al.
Pubblicazione: (2026)
Identifiable learning of dissipative dynamics
di: Zhu, Aiqing, et al.
Pubblicazione: (2025)
di: Zhu, Aiqing, et al.
Pubblicazione: (2025)
Mitigating distribution shift in machine learning-augmented hybrid simulation
di: Zhao, Jiaxi, et al.
Pubblicazione: (2024)
di: Zhao, Jiaxi, et al.
Pubblicazione: (2024)
Sequence-to-Sequence Models with Attention Mechanistically Map to the Architecture of Human Memory Search
di: Salvatore, Nikolaus, et al.
Pubblicazione: (2025)
di: Salvatore, Nikolaus, et al.
Pubblicazione: (2025)
Memory by Design: Probabilistic Sequence Layers
di: Dowling, Matthew, et al.
Pubblicazione: (2026)
di: Dowling, Matthew, et al.
Pubblicazione: (2026)
Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence Modeling
di: Zhang, Dehao, et al.
Pubblicazione: (2025)
di: Zhang, Dehao, et al.
Pubblicazione: (2025)
Mini-Sequence Transformer: Optimizing Intermediate Memory for Long Sequences Training
di: Luo, Cheng, et al.
Pubblicazione: (2024)
di: Luo, Cheng, et al.
Pubblicazione: (2024)
Terminally constrained flow-based generative models from an optimal control perspective
di: Gao, Weiguo, et al.
Pubblicazione: (2026)
di: Gao, Weiguo, et al.
Pubblicazione: (2026)
AutoChunk: Automated Activation Chunk for Memory-Efficient Long Sequence Inference
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024)
di: Zhao, Xuanlei, et al.
Pubblicazione: (2024)
Long-Sequence Memory with Temporal Kernels and Dense Hopfield Functionals
di: Farooq, Ahmed
Pubblicazione: (2025)
di: Farooq, Ahmed
Pubblicazione: (2025)
Language Models for Controllable DNA Sequence Design
di: Su, Xingyu, et al.
Pubblicazione: (2025)
di: Su, Xingyu, et al.
Pubblicazione: (2025)
Adaptive Substructure-Aware Expert Model for Molecular Property Prediction
di: Jiang, Tianyi, et al.
Pubblicazione: (2025)
di: Jiang, Tianyi, et al.
Pubblicazione: (2025)
Mamba-3: Improved Sequence Modeling using State Space Principles
di: Lahoti, Aakash, et al.
Pubblicazione: (2026)
di: Lahoti, Aakash, et al.
Pubblicazione: (2026)
RAM-Net: Expressive Linear Attention with Selectively Addressable Memory
di: Xiao, Kaicheng, et al.
Pubblicazione: (2026)
di: Xiao, Kaicheng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Approximation Rate of the Transformer Architecture for Sequence Modeling
di: Jiang, Haotian, et al.
Pubblicazione: (2023) -
StableSSM: Alleviating the Curse of Memory in State-space Models through Stable Reparameterization
di: Wang, Shida, et al.
Pubblicazione: (2023) -
The Effect of Depth on the Expressivity of Deep Linear State-Space Models
di: Bao, Zeyu, et al.
Pubblicazione: (2025) -
InfoFlow: A Framework for Multi-Layer Transformer Analysis
di: Yu, Penghao, et al.
Pubblicazione: (2026) -
Inverse Approximation Theory for Nonlinear Recurrent Neural Networks
di: Wang, Shida, et al.
Pubblicazione: (2023)