On the Expressiveness and Length Generalization of Selective State-Space Models on Regular Languages
Fuente:
arXiv
Saved in:
| Main Authors: | Terzić, Aleksandar, Hersche, Michael, Camposampiero, Giacomo, Hofmann, Thomas, Sebastian, Abu, Rahimi, Abbas |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Limits of Transformer Language Models on Learning to Compose Algorithms
by: Thomm, Jonathan, et al.
Published: (2024)
by: Thomm, Jonathan, et al.
Published: (2024)
Towards Learning Abductive Reasoning using VSA Distributed Representations
by: Camposampiero, Giacomo, et al.
Published: (2024)
by: Camposampiero, Giacomo, et al.
Published: (2024)
I-RAVEN-X: Benchmarking Generalization and Robustness of Analogical and Mathematical Reasoning in Large Language and Reasoning Models
by: Camposampiero, Giacomo, et al.
Published: (2025)
by: Camposampiero, Giacomo, et al.
Published: (2025)
Structured Sparse Transition Matrices to Enable State Tracking in State-Space Models
by: Terzić, Aleksandar, et al.
Published: (2025)
by: Terzić, Aleksandar, et al.
Published: (2025)
Can Large Reasoning Models do Analogical Reasoning under Perceptual Uncertainty?
by: Camposampiero, Giacomo, et al.
Published: (2025)
by: Camposampiero, Giacomo, et al.
Published: (2025)
Terminating Differentiable Tree Experts
by: Thomm, Jonathan, et al.
Published: (2024)
by: Thomm, Jonathan, et al.
Published: (2024)
Towards Learning to Reason: Comparing LLMs with Neuro-Symbolic on Arithmetic Relations in Abstract Reasoning
by: Hersche, Michael, et al.
Published: (2024)
by: Hersche, Michael, et al.
Published: (2024)
Scalable Evaluation and Neural Models for Compositional Generalization
by: Camposampiero, Giacomo, et al.
Published: (2025)
by: Camposampiero, Giacomo, et al.
Published: (2025)
Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models
by: Terzić, Aleksandar, et al.
Published: (2026)
by: Terzić, Aleksandar, et al.
Published: (2026)
Soft-Masked Diffusion Language Models
by: Hersche, Michael, et al.
Published: (2025)
by: Hersche, Michael, et al.
Published: (2025)
Thompson Sampling via Fine-Tuning of LLMs
by: Menet, Nicolas, et al.
Published: (2025)
by: Menet, Nicolas, et al.
Published: (2025)
Locally Coherent Parallel Decoding in Diffusion Language Models
by: Hersche, Michael, et al.
Published: (2026)
by: Hersche, Michael, et al.
Published: (2026)
On the Role of Noise in Factorizers for Disentangling Distributed Representations
by: Karunaratne, Geethan, et al.
Published: (2024)
by: Karunaratne, Geethan, et al.
Published: (2024)
Probabilistic Abduction for Visual Abstract Reasoning via Learning Rules in Vector-symbolic Architectures
by: Hersche, Michael, et al.
Published: (2024)
by: Hersche, Michael, et al.
Published: (2024)
A foundation model with multi-variate parallel attention to generate neuronal activity
by: Carzaniga, Francesco, et al.
Published: (2025)
by: Carzaniga, Francesco, et al.
Published: (2025)
Taipan: Efficient and Expressive State Space Language Models with Selective Attention
by: Van Nguyen, Chien, et al.
Published: (2024)
by: Van Nguyen, Chien, et al.
Published: (2024)
A Theoretical Analysis of Test-Driven Code Generation
by: Menet, Nicolas, et al.
Published: (2026)
by: Menet, Nicolas, et al.
Published: (2026)
Confidence Regularized Masked Language Modeling using Text Length
by: Ji, Seunghyun, et al.
Published: (2025)
by: Ji, Seunghyun, et al.
Published: (2025)
Retro-li: Small-Scale Retrieval Augmented Generation Supporting Noisy Similarity Searches and Domain Shift Generalization
by: Rashiti, Gentiana, et al.
Published: (2024)
by: Rashiti, Gentiana, et al.
Published: (2024)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
by: Pióro, Maciej, et al.
Published: (2024)
by: Pióro, Maciej, et al.
Published: (2024)
Sectoral Coupling in Linguistic State Space
by: Dumbrava, Sebastian
Published: (2025)
by: Dumbrava, Sebastian
Published: (2025)
Sessa: Selective State Space Attention
by: Horbatko, Liubomyr
Published: (2026)
by: Horbatko, Liubomyr
Published: (2026)
The Case for Cleaner Biosignals: High-fidelity Neural Compressor Enables Transfer from Cleaner iEEG to Noisier EEG
by: Carzaniga, Francesco Stefano, et al.
Published: (2025)
by: Carzaniga, Francesco Stefano, et al.
Published: (2025)
Length-MAX Tokenizer for Language Models
by: Dong, Dong, et al.
Published: (2025)
by: Dong, Dong, et al.
Published: (2025)
Context-Aware Initialization for Reducing Generative Path Length in Diffusion Language Models
by: Miao, Tongyuan, et al.
Published: (2025)
by: Miao, Tongyuan, et al.
Published: (2025)
LongSSM: On the Length Extension of State-space Models in Language Modelling
by: Wang, Shida
Published: (2024)
by: Wang, Shida
Published: (2024)
Projected Autoregression: Autoregressive Language Generation in Continuous State Space
by: Naparstek, Oshri
Published: (2026)
by: Naparstek, Oshri
Published: (2026)
Iterative Length-Regularized Direct Preference Optimization: A Case Study on Improving 7B Language Models to GPT-4 Level
by: Liu, Jie, et al.
Published: (2024)
by: Liu, Jie, et al.
Published: (2024)
Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language Models
by: Gao, Bo, et al.
Published: (2025)
by: Gao, Bo, et al.
Published: (2025)
Learning to (Learn at Test Time): RNNs with Expressive Hidden States
by: Sun, Yu, et al.
Published: (2024)
by: Sun, Yu, et al.
Published: (2024)
On the Optimal Reasoning Length for RL-Trained Language Models
by: Nohara, Daisuke, et al.
Published: (2026)
by: Nohara, Daisuke, et al.
Published: (2026)
Deep Learning Detection Method for Large Language Models-Generated Scientific Content
by: Alhijawi, Bushra, et al.
Published: (2024)
by: Alhijawi, Bushra, et al.
Published: (2024)
Large Language Models Badly Generalize across Option Length, Problem Types, and Irrelevant Noun Replacements
by: Zhao, Guangxiang, et al.
Published: (2025)
by: Zhao, Guangxiang, et al.
Published: (2025)
Factorizers for Distributed Sparse Block Codes
by: Hersche, Michael, et al.
Published: (2023)
by: Hersche, Michael, et al.
Published: (2023)
Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
by: Leng, Jiaqi, et al.
Published: (2025)
by: Leng, Jiaqi, et al.
Published: (2025)
Neural Diversity Regularizes Hallucinations in Language Models
by: Chakrabarti, Kushal, et al.
Published: (2025)
by: Chakrabarti, Kushal, et al.
Published: (2025)
Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results
by: Santilli, Andrea, et al.
Published: (2025)
by: Santilli, Andrea, et al.
Published: (2025)
SUV: Scalable Large Language Model Copyright Compliance with Regularized Selective Unlearning
by: Xu, Tianyang, et al.
Published: (2025)
by: Xu, Tianyang, et al.
Published: (2025)
The Role of Sparsity for Length Generalization in Transformers
by: Golowich, Noah, et al.
Published: (2025)
by: Golowich, Noah, et al.
Published: (2025)
Information Guided Regularization for Fine-tuning Language Models
by: Sharma, Mandar, et al.
Published: (2024)
by: Sharma, Mandar, et al.
Published: (2024)
Similar Items
-
Limits of Transformer Language Models on Learning to Compose Algorithms
by: Thomm, Jonathan, et al.
Published: (2024) -
Towards Learning Abductive Reasoning using VSA Distributed Representations
by: Camposampiero, Giacomo, et al.
Published: (2024) -
I-RAVEN-X: Benchmarking Generalization and Robustness of Analogical and Mathematical Reasoning in Large Language and Reasoning Models
by: Camposampiero, Giacomo, et al.
Published: (2025) -
Structured Sparse Transition Matrices to Enable State Tracking in State-Space Models
by: Terzić, Aleksandar, et al.
Published: (2025) -
Can Large Reasoning Models do Analogical Reasoning under Perceptual Uncertainty?
by: Camposampiero, Giacomo, et al.
Published: (2025)