A Provable Expressiveness Hierarchy in Hybrid Linear-Full Attention
Fuente:
arXiv
Salvato in:
| Autori principali: | Ye, Xiaowei, He, Xiaoyu, Liao, Chao, Wu, Chen, Lu, Pinyan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
di: Li, Xiaoyu, et al.
Pubblicazione: (2024)
di: Li, Xiaoyu, et al.
Pubblicazione: (2024)
On the Expressive Power and Limitations of Multi-Layer SSMs
di: Zubić, Nikola, et al.
Pubblicazione: (2026)
di: Zubić, Nikola, et al.
Pubblicazione: (2026)
Provably Overwhelming Transformer Models with Designed Inputs
di: Stambler, Lev, et al.
Pubblicazione: (2025)
di: Stambler, Lev, et al.
Pubblicazione: (2025)
Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent
di: Chen, Bo, et al.
Pubblicazione: (2025)
di: Chen, Bo, et al.
Pubblicazione: (2025)
RoPE Attention Can Be Trained in Almost Linear Time
di: Cao, Yang, et al.
Pubblicazione: (2024)
di: Cao, Yang, et al.
Pubblicazione: (2024)
On Computational Limits of FlowAR Models: Expressivity and Efficiency
di: Cao, Yang, et al.
Pubblicazione: (2025)
di: Cao, Yang, et al.
Pubblicazione: (2025)
On Fine-Grained I/O Complexity of Attention Backward Passes
di: Li, Xiaoyu, et al.
Pubblicazione: (2024)
di: Li, Xiaoyu, et al.
Pubblicazione: (2024)
On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis
di: Ke, Yekun, et al.
Pubblicazione: (2025)
di: Ke, Yekun, et al.
Pubblicazione: (2025)
On the Computational Capability of Graph Neural Networks: A Circuit Complexity Bound Perspective
di: Li, Xiaoyu, et al.
Pubblicazione: (2025)
di: Li, Xiaoyu, et al.
Pubblicazione: (2025)
The Computational Limits of State-Space Models and Mamba via the Lens of Circuit Complexity
di: Chen, Yifang, et al.
Pubblicazione: (2024)
di: Chen, Yifang, et al.
Pubblicazione: (2024)
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
di: Chen, Yifang, et al.
Pubblicazione: (2025)
di: Chen, Yifang, et al.
Pubblicazione: (2025)
Circuit Complexity Bounds for RoPE-based Transformer Architecture
di: Chen, Bo, et al.
Pubblicazione: (2024)
di: Chen, Bo, et al.
Pubblicazione: (2024)
Circuit Complexity Bounds for Visual Autoregressive Model
di: Ke, Yekun, et al.
Pubblicazione: (2025)
di: Ke, Yekun, et al.
Pubblicazione: (2025)
Modern Hopfield Networks Require Chain-of-Thought to Solve $\mathsf{NC}^1$-Hard Problems
di: Cao, Yang, et al.
Pubblicazione: (2024)
di: Cao, Yang, et al.
Pubblicazione: (2024)
Core Safety Values for Provably Corrigible Agents
di: Nayebi, Aran
Pubblicazione: (2025)
di: Nayebi, Aran
Pubblicazione: (2025)
A Quantitative Definition of Intelligence
di: Choi, Kang-Sin
Pubblicazione: (2026)
di: Choi, Kang-Sin
Pubblicazione: (2026)
A Theory of Learning with Autoregressive Chain of Thought
di: Joshi, Nirmit, et al.
Pubblicazione: (2025)
di: Joshi, Nirmit, et al.
Pubblicazione: (2025)
A Unified Approach for Maximizing Continuous DR-submodular Functions
di: Pedramfar, Mohammad, et al.
Pubblicazione: (2023)
di: Pedramfar, Mohammad, et al.
Pubblicazione: (2023)
Have Large Language Models Learned to Reason? A Characterization via 3-SAT Phase Transition
di: Hazra, Rishi, et al.
Pubblicazione: (2025)
di: Hazra, Rishi, et al.
Pubblicazione: (2025)
Learning to Think from Multiple Thinkers
di: Joshi, Nirmit, et al.
Pubblicazione: (2026)
di: Joshi, Nirmit, et al.
Pubblicazione: (2026)
How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers
di: Wang, Xiao
Pubblicazione: (2026)
di: Wang, Xiao
Pubblicazione: (2026)
Computational Limits of Low-Rank Adaptation (LoRA) Fine-Tuning for Transformer Models
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
Lossless Model Compression via Joint Low-Rank Factorization Optimization
di: Zhang, Boyang, et al.
Pubblicazione: (2024)
di: Zhang, Boyang, et al.
Pubblicazione: (2024)
Limitations on Accurate, Trusted, Human-level Reasoning
di: Panigrahy, Rina, et al.
Pubblicazione: (2025)
di: Panigrahy, Rina, et al.
Pubblicazione: (2025)
Looped ReLU MLPs May Be All You Need as Practical Programmable Computers
di: Liang, Yingyu, et al.
Pubblicazione: (2024)
di: Liang, Yingyu, et al.
Pubblicazione: (2024)
Mathematical Formalism for Memory Compression in Selective State Space Models
di: Bhat, Siddhanth
Pubblicazione: (2024)
di: Bhat, Siddhanth
Pubblicazione: (2024)
When Can We Solve the Weighted Low Rank Approximation Problem in Truly Subquadratic Time?
di: Li, Chenyang, et al.
Pubblicazione: (2025)
di: Li, Chenyang, et al.
Pubblicazione: (2025)
Mathematical Algorithm Design for Deep Learning under Societal and Judicial Constraints: The Algorithmic Transparency Requirement
di: Boche, Holger, et al.
Pubblicazione: (2024)
di: Boche, Holger, et al.
Pubblicazione: (2024)
A Measure-Theoretic Analysis of Reasoning: Structural Generalization and Approximation Limits
di: Zhang, Yuyang, et al.
Pubblicazione: (2026)
di: Zhang, Yuyang, et al.
Pubblicazione: (2026)
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
di: Svete, Anej, et al.
Pubblicazione: (2026)
di: Svete, Anej, et al.
Pubblicazione: (2026)
Theoretical limitations of multi-layer Transformer
di: Chen, Lijie, et al.
Pubblicazione: (2024)
di: Chen, Lijie, et al.
Pubblicazione: (2024)
A Complexity Map of Probabilistic Reasoning for Neurosymbolic Classification Techniques
di: Ledaguenel, Arthur, et al.
Pubblicazione: (2024)
di: Ledaguenel, Arthur, et al.
Pubblicazione: (2024)
Diffusion Language Models are Provably Optimal Parallel Samplers
di: Jiang, Haozhe, et al.
Pubblicazione: (2025)
di: Jiang, Haozhe, et al.
Pubblicazione: (2025)
Learning Tree Pattern Transformations
di: Neider, Daniel, et al.
Pubblicazione: (2024)
di: Neider, Daniel, et al.
Pubblicazione: (2024)
Nearest Neighbor CCP-Based Molecular Sequence Analysis
di: Ali, Sarwan, et al.
Pubblicazione: (2024)
di: Ali, Sarwan, et al.
Pubblicazione: (2024)
Reinforced Generation of Combinatorial Structures: Hardness of Approximation
di: Nagda, Ansh, et al.
Pubblicazione: (2025)
di: Nagda, Ansh, et al.
Pubblicazione: (2025)
Gateways to Tractability for Satisfiability in Pearl's Causal Hierarchy
di: Ganian, Robert, et al.
Pubblicazione: (2025)
di: Ganian, Robert, et al.
Pubblicazione: (2025)
Near-Optimal Learning and Planning in Separated Latent MDPs
di: Chen, Fan, et al.
Pubblicazione: (2024)
di: Chen, Fan, et al.
Pubblicazione: (2024)
Demystifying the unreasonable effectiveness of online alignment methods
di: Kang, Enoch Hyunwook
Pubblicazione: (2026)
di: Kang, Enoch Hyunwook
Pubblicazione: (2026)
The Alignment Trap: Complexity Barriers
di: Yao, Jasper
Pubblicazione: (2025)
di: Yao, Jasper
Pubblicazione: (2025)
Documenti analoghi
-
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
di: Li, Xiaoyu, et al.
Pubblicazione: (2024) -
On the Expressive Power and Limitations of Multi-Layer SSMs
di: Zubić, Nikola, et al.
Pubblicazione: (2026) -
Provably Overwhelming Transformer Models with Designed Inputs
di: Stambler, Lev, et al.
Pubblicazione: (2025) -
Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent
di: Chen, Bo, et al.
Pubblicazione: (2025) -
RoPE Attention Can Be Trained in Almost Linear Time
di: Cao, Yang, et al.
Pubblicazione: (2024)