LongSSM: On the Length Extension of State-space Models in Language Modelling
Fuente:
arXiv
Saved in:
| Main Author: | Wang, Shida |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StableSSM: Alleviating the Curse of Memory in State-space Models through Stable Reparameterization
by: Wang, Shida, et al.
Published: (2023)
by: Wang, Shida, et al.
Published: (2023)
Inverse Approximation Theory for Nonlinear Recurrent Neural Networks
by: Wang, Shida, et al.
Published: (2023)
by: Wang, Shida, et al.
Published: (2023)
Zamba: A Compact 7B SSM Hybrid Model
by: Glorioso, Paolo, et al.
Published: (2024)
by: Glorioso, Paolo, et al.
Published: (2024)
On the Expressiveness and Length Generalization of Selective State-Space Models on Regular Languages
by: Terzić, Aleksandar, et al.
Published: (2024)
by: Terzić, Aleksandar, et al.
Published: (2024)
Local Normalization Distortion and the Thermodynamic Formalism of Decoding Strategies for Large Language Models
by: Kempton, Tom, et al.
Published: (2025)
by: Kempton, Tom, et al.
Published: (2025)
Length-MAX Tokenizer for Language Models
by: Dong, Dong, et al.
Published: (2025)
by: Dong, Dong, et al.
Published: (2025)
Language models as master equation solvers
by: Liu, Chuanbo, et al.
Published: (2023)
by: Liu, Chuanbo, et al.
Published: (2023)
True Zero-Shot Inference of Dynamical Systems Preserving Long-Term Statistics
by: Hemmer, Christoph Jürgen, et al.
Published: (2025)
by: Hemmer, Christoph Jürgen, et al.
Published: (2025)
Unified Long-Term Time-Series Forecasting Benchmark
by: Cyranka, Jacek, et al.
Published: (2023)
by: Cyranka, Jacek, et al.
Published: (2023)
On the Optimal Reasoning Length for RL-Trained Language Models
by: Nohara, Daisuke, et al.
Published: (2026)
by: Nohara, Daisuke, et al.
Published: (2026)
The Recurrent Sticky Hierarchical Dirichlet Process Hidden Markov Model
by: Słupiński, Mikołaj, et al.
Published: (2024)
by: Słupiński, Mikołaj, et al.
Published: (2024)
Dynamical Systems Theory Behind a Hierarchical Reasoning Model
by: Es'kin, Vasiliy A., et al.
Published: (2026)
by: Es'kin, Vasiliy A., et al.
Published: (2026)
Editing Personality for Large Language Models
by: Mao, Shengyu, et al.
Published: (2023)
by: Mao, Shengyu, et al.
Published: (2023)
Confidence Regularized Masked Language Modeling using Text Length
by: Ji, Seunghyun, et al.
Published: (2025)
by: Ji, Seunghyun, et al.
Published: (2025)
Large-Step Training Dynamics of a Two-Factor Linear Transformer Model
by: Balasubramanian, Krishnakumar
Published: (2026)
by: Balasubramanian, Krishnakumar
Published: (2026)
Memory in Plain Sight: Surveying the Uncanny Resemblances of Associative Memories and Diffusion Models
by: Hoover, Benjamin, et al.
Published: (2023)
by: Hoover, Benjamin, et al.
Published: (2023)
Position: Why a Dynamical Systems Perspective is Needed to Advance Time Series Modeling
by: Durstewitz, Daniel, et al.
Published: (2026)
by: Durstewitz, Daniel, et al.
Published: (2026)
A Minimal Bifurcation Model of Load Imbalance in a Softmax Mixture-of-Experts Router
by: Kiselev, O. M.
Published: (2026)
by: Kiselev, O. M.
Published: (2026)
In-Context Environments Induce Evaluation-Awareness in Language Models
by: Chaudhary, Maheep
Published: (2026)
by: Chaudhary, Maheep
Published: (2026)
Multiscale Byte Language Models -- A Hierarchical Architecture for Causal Million-Length Sequence Modeling
by: Egli, Eric, et al.
Published: (2025)
by: Egli, Eric, et al.
Published: (2025)
A Hybrid Approach of Transfer Learning and Physics-Informed Modeling: Improving Dissolved Oxygen Concentration Prediction in an Industrial Wastewater Treatment Plant
by: Koksal, Ece S., et al.
Published: (2024)
by: Koksal, Ece S., et al.
Published: (2024)
YaRN: Efficient Context Window Extension of Large Language Models
by: Peng, Bowen, et al.
Published: (2023)
by: Peng, Bowen, et al.
Published: (2023)
Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning
by: Sarkar, Bidipta, et al.
Published: (2025)
by: Sarkar, Bidipta, et al.
Published: (2025)
Uncovering the Computational Roles of Nonlinearity in Sequence Modeling Using Almost-Linear RNNs
by: Brenner, Manuel, et al.
Published: (2025)
by: Brenner, Manuel, et al.
Published: (2025)
A Survey on Large Language Model-empowered Autonomous Driving
by: Zhu, Yuxuan, et al.
Published: (2024)
by: Zhu, Yuxuan, et al.
Published: (2024)
Context-Aware Initialization for Reducing Generative Path Length in Diffusion Language Models
by: Miao, Tongyuan, et al.
Published: (2025)
by: Miao, Tongyuan, et al.
Published: (2025)
Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language Models
by: Gao, Bo, et al.
Published: (2025)
by: Gao, Bo, et al.
Published: (2025)
Fine-tuning Smaller Language Models for Question Answering over Financial Documents
by: Phogat, Karmvir Singh, et al.
Published: (2024)
by: Phogat, Karmvir Singh, et al.
Published: (2024)
Iteration of Thought: Leveraging Inner Dialogue for Autonomous Large Language Model Reasoning
by: Radha, Santosh Kumar, et al.
Published: (2024)
by: Radha, Santosh Kumar, et al.
Published: (2024)
Dipper: Diversity in Prompts for Producing Large Language Model Ensembles in Reasoning tasks
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
Multi-Objective Reinforcement Learning for Large Language Model Optimization: Visionary Perspective
by: Kong, Lingxiao, et al.
Published: (2025)
by: Kong, Lingxiao, et al.
Published: (2025)
Learning Chaotic Systems and Long-Term Predictions with Neural Jump ODEs
by: Krach, Florian, et al.
Published: (2024)
by: Krach, Florian, et al.
Published: (2024)
Hybrid Energy-Based Models for Physical AI: Provably Stable Identification of Port-Hamiltonian Dynamics
by: Betteti, Simone, et al.
Published: (2026)
by: Betteti, Simone, et al.
Published: (2026)
Language Models as Efficient Reward Function Searchers for Custom-Environment Multi-Objective Reinforcement
by: Xie, Guanwen, et al.
Published: (2024)
by: Xie, Guanwen, et al.
Published: (2024)
AdjointDEIS: Efficient Gradients for Diffusion Models
by: Blasingame, Zander W., et al.
Published: (2024)
by: Blasingame, Zander W., et al.
Published: (2024)
Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
by: Leng, Jiaqi, et al.
Published: (2025)
by: Leng, Jiaqi, et al.
Published: (2025)
More Thinking, More Bias: Length-Driven Position Bias in Reasoning Models
by: Wang, Xiao
Published: (2026)
by: Wang, Xiao
Published: (2026)
Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models
by: Bozdag, Nimet Beyza, et al.
Published: (2025)
by: Bozdag, Nimet Beyza, et al.
Published: (2025)
Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results
by: Santilli, Andrea, et al.
Published: (2025)
by: Santilli, Andrea, et al.
Published: (2025)
Effects of Prompt Length on Domain-specific Tasks for Large Language Models
by: Liu, Qibang, et al.
Published: (2025)
by: Liu, Qibang, et al.
Published: (2025)
Similar Items
-
StableSSM: Alleviating the Curse of Memory in State-space Models through Stable Reparameterization
by: Wang, Shida, et al.
Published: (2023) -
Inverse Approximation Theory for Nonlinear Recurrent Neural Networks
by: Wang, Shida, et al.
Published: (2023) -
Zamba: A Compact 7B SSM Hybrid Model
by: Glorioso, Paolo, et al.
Published: (2024) -
On the Expressiveness and Length Generalization of Selective State-Space Models on Regular Languages
by: Terzić, Aleksandar, et al.
Published: (2024) -
Local Normalization Distortion and the Thermodynamic Formalism of Decoding Strategies for Large Language Models
by: Kempton, Tom, et al.
Published: (2025)