The Expressive Limits of Diagonal SSMs for State-Tracking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shakerinava, Mehran, Khavari, Behnoush, Ravanbakhsh, Siamak, Chandar, Sarath |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Parity Requires Unified Input Dependence and Negative Eigenvalues in SSMs
von: Khavari, Behnoush, et al.
Veröffentlicht: (2025)
von: Khavari, Behnoush, et al.
Veröffentlicht: (2025)
Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
von: Shakerinava, Mehran, et al.
Veröffentlicht: (2025)
von: Shakerinava, Mehran, et al.
Veröffentlicht: (2025)
Weight-Sharing Regularization
von: Shakerinava, Mehran, et al.
Veröffentlicht: (2023)
von: Shakerinava, Mehran, et al.
Veröffentlicht: (2023)
The Role of Symmetry in Optimizing Overparameterized Networks
von: Sareen, Kusha, et al.
Veröffentlicht: (2026)
von: Sareen, Kusha, et al.
Veröffentlicht: (2026)
On the Expressive Power and Limitations of Multi-Layer SSMs
von: Zubić, Nikola, et al.
Veröffentlicht: (2026)
von: Zubić, Nikola, et al.
Veröffentlicht: (2026)
Learning to Reach Goals via Diffusion
von: Jain, Vineet, et al.
Veröffentlicht: (2023)
von: Jain, Vineet, et al.
Veröffentlicht: (2023)
Multi-Armed Sampling Problem and the End of Exploration
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2025)
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2025)
Symmetry Breaking and Equivariant Neural Networks
von: Kaba, Sékou-Oumar, et al.
Veröffentlicht: (2023)
von: Kaba, Sékou-Oumar, et al.
Veröffentlicht: (2023)
Scaling Laws and Symmetry, Evidence from Neural Force Fields
von: Ngo, Khang, et al.
Veröffentlicht: (2025)
von: Ngo, Khang, et al.
Veröffentlicht: (2025)
On the Identifiability of Causal Abstractions
von: Li, Xiusi, et al.
Veröffentlicht: (2025)
von: Li, Xiusi, et al.
Veröffentlicht: (2025)
Sampling from Energy-based Policies using Diffusion
von: Jain, Vineet, et al.
Veröffentlicht: (2024)
von: Jain, Vineet, et al.
Veröffentlicht: (2024)
Faithfulness Measurable Masked Language Models
von: Madsen, Andreas, et al.
Veröffentlicht: (2023)
von: Madsen, Andreas, et al.
Veröffentlicht: (2023)
On Diffusion Modeling for Anomaly Detection
von: Livernoche, Victor, et al.
Veröffentlicht: (2023)
von: Livernoche, Victor, et al.
Veröffentlicht: (2023)
Diffusion Tree Sampling: Scalable inference-time alignment of diffusion models
von: Jain, Vineet, et al.
Veröffentlicht: (2025)
von: Jain, Vineet, et al.
Veröffentlicht: (2025)
Are self-explanations from Large Language Models faithful?
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
Exploring Quantization for Efficient Pre-Training of Transformer Language Models
von: Chitsaz, Kamran, et al.
Veröffentlicht: (2024)
von: Chitsaz, Kamran, et al.
Veröffentlicht: (2024)
GG-SSMs: Graph-Generating State Space Models
von: Zubić, Nikola, et al.
Veröffentlicht: (2024)
von: Zubić, Nikola, et al.
Veröffentlicht: (2024)
Mastering Memory Tasks with World Models
von: Samsami, Mohammad Reza, et al.
Veröffentlicht: (2024)
von: Samsami, Mohammad Reza, et al.
Veröffentlicht: (2024)
Neural Coherence : Find higher performance to out-of-distribution tasks from few samples
von: Guiroy, Simon, et al.
Veröffentlicht: (2025)
von: Guiroy, Simon, et al.
Veröffentlicht: (2025)
Intelligent Switching for Reset-Free RL
von: Patil, Darshan, et al.
Veröffentlicht: (2024)
von: Patil, Darshan, et al.
Veröffentlicht: (2024)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
Lookbehind-SAM: k steps back, 1 step forward
von: Mordido, Gonçalo, et al.
Veröffentlicht: (2023)
von: Mordido, Gonçalo, et al.
Veröffentlicht: (2023)
Energy Loss Functions for Physical Systems
von: Kaba, Sékou-Oumar, et al.
Veröffentlicht: (2025)
von: Kaba, Sékou-Oumar, et al.
Veröffentlicht: (2025)
Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models
von: Nilaksh, et al.
Veröffentlicht: (2026)
von: Nilaksh, et al.
Veröffentlicht: (2026)
CoPeP: Benchmarking Continual Pretraining for Protein Language Models
von: Patil, Darshan, et al.
Veröffentlicht: (2026)
von: Patil, Darshan, et al.
Veröffentlicht: (2026)
Sub-goal Distillation: A Method to Improve Small Language Agents
von: Hashemzadeh, Maryam, et al.
Veröffentlicht: (2024)
von: Hashemzadeh, Maryam, et al.
Veröffentlicht: (2024)
Manifold Metric: A Loss Landscape Approach for Predicting Model Performance
von: Malviya, Pranshu, et al.
Veröffentlicht: (2024)
von: Malviya, Pranshu, et al.
Veröffentlicht: (2024)
Squeezing More from the Stream : Learning Representation Online for Streaming Reinforcement Learning
von: Nilaksh, et al.
Veröffentlicht: (2026)
von: Nilaksh, et al.
Veröffentlicht: (2026)
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
Towards Practical Tool Usage for Continually Learning LLMs
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
Inverting Data Transformations via Diffusion Sampling
von: Kim, Jinwoo, et al.
Veröffentlicht: (2026)
von: Kim, Jinwoo, et al.
Veröffentlicht: (2026)
Lag Operator SSMs: A Geometric Framework for Structured State Space Modeling
von: Tomonaga, Sutashu, et al.
Veröffentlicht: (2025)
von: Tomonaga, Sutashu, et al.
Veröffentlicht: (2025)
Dialectics of Alignment: Harnessing Unsafe Knowledge for Dynamic Safety Routing
von: Hashemzadeh, Maryam, et al.
Veröffentlicht: (2026)
von: Hashemzadeh, Maryam, et al.
Veröffentlicht: (2026)
NovoMolGen: Rethinking Molecular Language Model Pretraining
von: Chitsaz, Kamran, et al.
Veröffentlicht: (2025)
von: Chitsaz, Kamran, et al.
Veröffentlicht: (2025)
Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit Conservatism
von: Ni, Tianwei, et al.
Veröffentlicht: (2025)
von: Ni, Tianwei, et al.
Veröffentlicht: (2025)
Interpretability Needs a New Paradigm
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
Why Don't Prompt-Based Fairness Metrics Correlate?
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2024)
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2024)
Should We Attend More or Less? Modulating Attention for Fairness
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2023)
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2023)
GRPO-$λ$: Credit Assignment improves LLM Reasoning
von: Parthasarathi, Prasanna, et al.
Veröffentlicht: (2025)
von: Parthasarathi, Prasanna, et al.
Veröffentlicht: (2025)
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
von: Dao, Tri, et al.
Veröffentlicht: (2024)
von: Dao, Tri, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Parity Requires Unified Input Dependence and Negative Eigenvalues in SSMs
von: Khavari, Behnoush, et al.
Veröffentlicht: (2025) -
Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
von: Shakerinava, Mehran, et al.
Veröffentlicht: (2025) -
Weight-Sharing Regularization
von: Shakerinava, Mehran, et al.
Veröffentlicht: (2023) -
The Role of Symmetry in Optimizing Overparameterized Networks
von: Sareen, Kusha, et al.
Veröffentlicht: (2026) -
On the Expressive Power and Limitations of Multi-Layer SSMs
von: Zubić, Nikola, et al.
Veröffentlicht: (2026)