Numerical Investigation of Sequence Modeling Theory using Controllable Memory Functions
Fuente:
arXiv
Guardado en:
| Autores principales: | Jiang, Haotian, Bao, Zeyu, Wang, Shida, Li, Qianxiao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Approximation Rate of the Transformer Architecture for Sequence Modeling
por: Jiang, Haotian, et al.
Publicado: (2023)
por: Jiang, Haotian, et al.
Publicado: (2023)
StableSSM: Alleviating the Curse of Memory in State-space Models through Stable Reparameterization
por: Wang, Shida, et al.
Publicado: (2023)
por: Wang, Shida, et al.
Publicado: (2023)
The Effect of Depth on the Expressivity of Deep Linear State-Space Models
por: Bao, Zeyu, et al.
Publicado: (2025)
por: Bao, Zeyu, et al.
Publicado: (2025)
InfoFlow: A Framework for Multi-Layer Transformer Analysis
por: Yu, Penghao, et al.
Publicado: (2026)
por: Yu, Penghao, et al.
Publicado: (2026)
Inverse Approximation Theory for Nonlinear Recurrent Neural Networks
por: Wang, Shida, et al.
Publicado: (2023)
por: Wang, Shida, et al.
Publicado: (2023)
The Effect of Attention Head Count on Transformer Approximation
por: Yu, Penghao, et al.
Publicado: (2025)
por: Yu, Penghao, et al.
Publicado: (2025)
Accelerating Legacy Numerical Solvers by Non-intrusive Gradient-based Meta-solving
por: Arisaka, Sohei, et al.
Publicado: (2024)
por: Arisaka, Sohei, et al.
Publicado: (2024)
From Generalization Analysis to Optimization Designs for State Space Models
por: Liu, Fusheng, et al.
Publicado: (2024)
por: Liu, Fusheng, et al.
Publicado: (2024)
Autocorrelation Matters: Understanding the Role of Initialization Schemes for State Space Models
por: Liu, Fusheng, et al.
Publicado: (2024)
por: Liu, Fusheng, et al.
Publicado: (2024)
Allocation of Parameters in Transformers
por: Yu, Ruoxi, et al.
Publicado: (2025)
por: Yu, Ruoxi, et al.
Publicado: (2025)
Learning task-specific predictive models for scientific computing
por: Yin, Jianyuan, et al.
Publicado: (2025)
por: Yin, Jianyuan, et al.
Publicado: (2025)
Unifying back-propagation and forward-forward algorithms through model predictive control
por: Ren, Lianhai, et al.
Publicado: (2024)
por: Ren, Lianhai, et al.
Publicado: (2024)
LongSSM: On the Length Extension of State-space Models in Language Modelling
por: Wang, Shida
Publicado: (2024)
por: Wang, Shida
Publicado: (2024)
DynGMA: a robust approach for learning stochastic differential equations from data
por: Zhu, Aiqing, et al.
Publicado: (2024)
por: Zhu, Aiqing, et al.
Publicado: (2024)
Learning Macroscopic Dynamics from Partial Microscopic Observations
por: Chen, Mengyi, et al.
Publicado: (2024)
por: Chen, Mengyi, et al.
Publicado: (2024)
Continuity-Preserving Convolutional Autoencoders for Learning Continuous Latent Dynamical Models from Images
por: Zhu, Aiqing, et al.
Publicado: (2025)
por: Zhu, Aiqing, et al.
Publicado: (2025)
LongVQ: Long Sequence Modeling with Vector Quantization on Structured Memory
por: Liu, Zicheng, et al.
Publicado: (2024)
por: Liu, Zicheng, et al.
Publicado: (2024)
Learning Permutation-invariant Macroscopic Dynamics
por: Han, Zhichao, et al.
Publicado: (2026)
por: Han, Zhichao, et al.
Publicado: (2026)
Machine Unlearning under Retain-Forget Entanglement
por: Cheng, Jingpu, et al.
Publicado: (2026)
por: Cheng, Jingpu, et al.
Publicado: (2026)
On the Curse of Memory in Recurrent Neural Networks: Approximation and Optimization Analysis
por: Li, Zhong, et al.
Publicado: (2020)
por: Li, Zhong, et al.
Publicado: (2020)
On Hypothesis Transfer Learning of Functional Linear Models
por: Lin, Haotian, et al.
Publicado: (2022)
por: Lin, Haotian, et al.
Publicado: (2022)
SMR: State Memory Replay for Long Sequence Modeling
por: Qi, Biqing, et al.
Publicado: (2024)
por: Qi, Biqing, et al.
Publicado: (2024)
Multi-Modal Representation Learning for Molecular Property Prediction: Sequence, Graph, Geometry
por: Wang, Zeyu, et al.
Publicado: (2024)
por: Wang, Zeyu, et al.
Publicado: (2024)
MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration
por: Ren, Lianhai, et al.
Publicado: (2026)
por: Ren, Lianhai, et al.
Publicado: (2026)
Scalable learning of macroscopic stochastic dynamics
por: Chen, Mengyi, et al.
Publicado: (2025)
por: Chen, Mengyi, et al.
Publicado: (2025)
A unified framework for establishing the universal approximation of transformer-type architectures
por: Cheng, Jingpu, et al.
Publicado: (2025)
por: Cheng, Jingpu, et al.
Publicado: (2025)
Deep learning and the rate of approximation by flows
por: Cheng, Jingpu, et al.
Publicado: (2026)
por: Cheng, Jingpu, et al.
Publicado: (2026)
Identifiable learning of dissipative dynamics
por: Zhu, Aiqing, et al.
Publicado: (2025)
por: Zhu, Aiqing, et al.
Publicado: (2025)
Mitigating distribution shift in machine learning-augmented hybrid simulation
por: Zhao, Jiaxi, et al.
Publicado: (2024)
por: Zhao, Jiaxi, et al.
Publicado: (2024)
Sequence-to-Sequence Models with Attention Mechanistically Map to the Architecture of Human Memory Search
por: Salvatore, Nikolaus, et al.
Publicado: (2025)
por: Salvatore, Nikolaus, et al.
Publicado: (2025)
Memory by Design: Probabilistic Sequence Layers
por: Dowling, Matthew, et al.
Publicado: (2026)
por: Dowling, Matthew, et al.
Publicado: (2026)
Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence Modeling
por: Zhang, Dehao, et al.
Publicado: (2025)
por: Zhang, Dehao, et al.
Publicado: (2025)
Mini-Sequence Transformer: Optimizing Intermediate Memory for Long Sequences Training
por: Luo, Cheng, et al.
Publicado: (2024)
por: Luo, Cheng, et al.
Publicado: (2024)
Terminally constrained flow-based generative models from an optimal control perspective
por: Gao, Weiguo, et al.
Publicado: (2026)
por: Gao, Weiguo, et al.
Publicado: (2026)
AutoChunk: Automated Activation Chunk for Memory-Efficient Long Sequence Inference
por: Zhao, Xuanlei, et al.
Publicado: (2024)
por: Zhao, Xuanlei, et al.
Publicado: (2024)
Long-Sequence Memory with Temporal Kernels and Dense Hopfield Functionals
por: Farooq, Ahmed
Publicado: (2025)
por: Farooq, Ahmed
Publicado: (2025)
Language Models for Controllable DNA Sequence Design
por: Su, Xingyu, et al.
Publicado: (2025)
por: Su, Xingyu, et al.
Publicado: (2025)
Adaptive Substructure-Aware Expert Model for Molecular Property Prediction
por: Jiang, Tianyi, et al.
Publicado: (2025)
por: Jiang, Tianyi, et al.
Publicado: (2025)
Mamba-3: Improved Sequence Modeling using State Space Principles
por: Lahoti, Aakash, et al.
Publicado: (2026)
por: Lahoti, Aakash, et al.
Publicado: (2026)
RAM-Net: Expressive Linear Attention with Selectively Addressable Memory
por: Xiao, Kaicheng, et al.
Publicado: (2026)
por: Xiao, Kaicheng, et al.
Publicado: (2026)
Ejemplares similares
-
Approximation Rate of the Transformer Architecture for Sequence Modeling
por: Jiang, Haotian, et al.
Publicado: (2023) -
StableSSM: Alleviating the Curse of Memory in State-space Models through Stable Reparameterization
por: Wang, Shida, et al.
Publicado: (2023) -
The Effect of Depth on the Expressivity of Deep Linear State-Space Models
por: Bao, Zeyu, et al.
Publicado: (2025) -
InfoFlow: A Framework for Multi-Layer Transformer Analysis
por: Yu, Penghao, et al.
Publicado: (2026) -
Inverse Approximation Theory for Nonlinear Recurrent Neural Networks
por: Wang, Shida, et al.
Publicado: (2023)