Mathematical Formalism for Memory Compression in Selective State Space Models
Fuente:
arXiv
Guardado en:
| Autor principal: | Bhat, Siddhanth |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
por: Chen, Yifang, et al.
Publicado: (2025)
por: Chen, Yifang, et al.
Publicado: (2025)
Lossless Model Compression via Joint Low-Rank Factorization Optimization
por: Zhang, Boyang, et al.
Publicado: (2024)
por: Zhang, Boyang, et al.
Publicado: (2024)
The Computational Limits of State-Space Models and Mamba via the Lens of Circuit Complexity
por: Chen, Yifang, et al.
Publicado: (2024)
por: Chen, Yifang, et al.
Publicado: (2024)
The Computational Complexity of Satisfiability in State Space Models
por: Alsmann, Eric, et al.
Publicado: (2025)
por: Alsmann, Eric, et al.
Publicado: (2025)
Mathematical Algorithm Design for Deep Learning under Societal and Judicial Constraints: The Algorithmic Transparency Requirement
por: Boche, Holger, et al.
Publicado: (2024)
por: Boche, Holger, et al.
Publicado: (2024)
How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers
por: Wang, Xiao
Publicado: (2026)
por: Wang, Xiao
Publicado: (2026)
Transformer Encoder Satisfiability: Complexity and Impact on Formal Reasoning
por: Sälzer, Marco, et al.
Publicado: (2024)
por: Sälzer, Marco, et al.
Publicado: (2024)
Provably Overwhelming Transformer Models with Designed Inputs
por: Stambler, Lev, et al.
Publicado: (2025)
por: Stambler, Lev, et al.
Publicado: (2025)
Computational Limits of Low-Rank Adaptation (LoRA) Fine-Tuning for Transformer Models
por: Hu, Jerry Yao-Chieh, et al.
Publicado: (2024)
por: Hu, Jerry Yao-Chieh, et al.
Publicado: (2024)
Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent
por: Chen, Bo, et al.
Publicado: (2025)
por: Chen, Bo, et al.
Publicado: (2025)
Have Large Language Models Learned to Reason? A Characterization via 3-SAT Phase Transition
por: Hazra, Rishi, et al.
Publicado: (2025)
por: Hazra, Rishi, et al.
Publicado: (2025)
Compression Barriers for Autoregressive Transformers
por: Haris, Themistoklis, et al.
Publicado: (2025)
por: Haris, Themistoklis, et al.
Publicado: (2025)
Looped ReLU MLPs May Be All You Need as Practical Programmable Computers
por: Liang, Yingyu, et al.
Publicado: (2024)
por: Liang, Yingyu, et al.
Publicado: (2024)
Learning to Think from Multiple Thinkers
por: Joshi, Nirmit, et al.
Publicado: (2026)
por: Joshi, Nirmit, et al.
Publicado: (2026)
A Theory of Learning with Autoregressive Chain of Thought
por: Joshi, Nirmit, et al.
Publicado: (2025)
por: Joshi, Nirmit, et al.
Publicado: (2025)
A Provable Expressiveness Hierarchy in Hybrid Linear-Full Attention
por: Ye, Xiaowei, et al.
Publicado: (2026)
por: Ye, Xiaowei, et al.
Publicado: (2026)
On the Expressive Power and Limitations of Multi-Layer SSMs
por: Zubić, Nikola, et al.
Publicado: (2026)
por: Zubić, Nikola, et al.
Publicado: (2026)
Limitations on Accurate, Trusted, Human-level Reasoning
por: Panigrahy, Rina, et al.
Publicado: (2025)
por: Panigrahy, Rina, et al.
Publicado: (2025)
When Can We Solve the Weighted Low Rank Approximation Problem in Truly Subquadratic Time?
por: Li, Chenyang, et al.
Publicado: (2025)
por: Li, Chenyang, et al.
Publicado: (2025)
A Unified Approach for Maximizing Continuous DR-submodular Functions
por: Pedramfar, Mohammad, et al.
Publicado: (2023)
por: Pedramfar, Mohammad, et al.
Publicado: (2023)
A Quantitative Definition of Intelligence
por: Choi, Kang-Sin
Publicado: (2026)
por: Choi, Kang-Sin
Publicado: (2026)
On the Computational Capability of Graph Neural Networks: A Circuit Complexity Bound Perspective
por: Li, Xiaoyu, et al.
Publicado: (2025)
por: Li, Xiaoyu, et al.
Publicado: (2025)
Circuit Complexity Bounds for Visual Autoregressive Model
por: Ke, Yekun, et al.
Publicado: (2025)
por: Ke, Yekun, et al.
Publicado: (2025)
Emotion Classification in Short English Texts using Deep Learning Techniques
por: Bhat, Siddhanth
Publicado: (2024)
por: Bhat, Siddhanth
Publicado: (2024)
NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes
por: Fan, Lizhou, et al.
Publicado: (2023)
por: Fan, Lizhou, et al.
Publicado: (2023)
Learning Tree Pattern Transformations
por: Neider, Daniel, et al.
Publicado: (2024)
por: Neider, Daniel, et al.
Publicado: (2024)
Nearest Neighbor CCP-Based Molecular Sequence Analysis
por: Ali, Sarwan, et al.
Publicado: (2024)
por: Ali, Sarwan, et al.
Publicado: (2024)
Reinforced Generation of Combinatorial Structures: Hardness of Approximation
por: Nagda, Ansh, et al.
Publicado: (2025)
por: Nagda, Ansh, et al.
Publicado: (2025)
Data-driven Power Loss Identification through Physics-Based Thermal Model Backpropagation
por: Scarpa, Mattia, et al.
Publicado: (2025)
por: Scarpa, Mattia, et al.
Publicado: (2025)
RoPE Attention Can Be Trained in Almost Linear Time
por: Cao, Yang, et al.
Publicado: (2024)
por: Cao, Yang, et al.
Publicado: (2024)
Modern Hopfield Networks Require Chain-of-Thought to Solve $\mathsf{NC}^1$-Hard Problems
por: Cao, Yang, et al.
Publicado: (2024)
por: Cao, Yang, et al.
Publicado: (2024)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
por: Li, Xiaoyu, et al.
Publicado: (2024)
por: Li, Xiaoyu, et al.
Publicado: (2024)
A Complexity Map of Probabilistic Reasoning for Neurosymbolic Classification Techniques
por: Ledaguenel, Arthur, et al.
Publicado: (2024)
por: Ledaguenel, Arthur, et al.
Publicado: (2024)
Circuit Complexity Bounds for RoPE-based Transformer Architecture
por: Chen, Bo, et al.
Publicado: (2024)
por: Chen, Bo, et al.
Publicado: (2024)
On Fine-Grained I/O Complexity of Attention Backward Passes
por: Li, Xiaoyu, et al.
Publicado: (2024)
por: Li, Xiaoyu, et al.
Publicado: (2024)
Unlocking the Theory Behind Scaling 1-Bit Neural Networks
por: Daliri, Majid, et al.
Publicado: (2024)
por: Daliri, Majid, et al.
Publicado: (2024)
A Measure-Theoretic Analysis of Reasoning: Structural Generalization and Approximation Limits
por: Zhang, Yuyang, et al.
Publicado: (2026)
por: Zhang, Yuyang, et al.
Publicado: (2026)
The Alignment Trap: Complexity Barriers
por: Yao, Jasper
Publicado: (2025)
por: Yao, Jasper
Publicado: (2025)
Demystifying the unreasonable effectiveness of online alignment methods
por: Kang, Enoch Hyunwook
Publicado: (2026)
por: Kang, Enoch Hyunwook
Publicado: (2026)
Neural Algorithmic Reasoning for Hypergraphs with Looped Transformers
por: Huang, Zekai, et al.
Publicado: (2025)
por: Huang, Zekai, et al.
Publicado: (2025)
Ejemplares similares
-
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
por: Chen, Yifang, et al.
Publicado: (2025) -
Lossless Model Compression via Joint Low-Rank Factorization Optimization
por: Zhang, Boyang, et al.
Publicado: (2024) -
The Computational Limits of State-Space Models and Mamba via the Lens of Circuit Complexity
por: Chen, Yifang, et al.
Publicado: (2024) -
The Computational Complexity of Satisfiability in State Space Models
por: Alsmann, Eric, et al.
Publicado: (2025) -
Mathematical Algorithm Design for Deep Learning under Societal and Judicial Constraints: The Algorithmic Transparency Requirement
por: Boche, Holger, et al.
Publicado: (2024)