Looped ReLU MLPs May Be All You Need as Practical Programmable Computers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Yingyu, Sha, Zhizhou, Shi, Zhenmei, Song, Zhao, Zhou, Yufa |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis
von: Ke, Yekun, et al.
Veröffentlicht: (2025)
von: Ke, Yekun, et al.
Veröffentlicht: (2025)
Multi-Layer Transformers Gradient Can be Approximated in Almost Linear Time
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
On Computational Limits of FlowAR Models: Expressivity and Efficiency
von: Cao, Yang, et al.
Veröffentlicht: (2025)
von: Cao, Yang, et al.
Veröffentlicht: (2025)
Neural Algorithmic Reasoning for Hypergraphs with Looped Transformers
von: Huang, Zekai, et al.
Veröffentlicht: (2025)
von: Huang, Zekai, et al.
Veröffentlicht: (2025)
When Can We Solve the Weighted Low Rank Approximation Problem in Truly Subquadratic Time?
von: Li, Chenyang, et al.
Veröffentlicht: (2025)
von: Li, Chenyang, et al.
Veröffentlicht: (2025)
The Computational Limits of State-Space Models and Mamba via the Lens of Circuit Complexity
von: Chen, Yifang, et al.
Veröffentlicht: (2024)
von: Chen, Yifang, et al.
Veröffentlicht: (2024)
On the Computational Capability of Graph Neural Networks: A Circuit Complexity Bound Perspective
von: Li, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Li, Xiaoyu, et al.
Veröffentlicht: (2025)
RoPE Attention Can Be Trained in Almost Linear Time
von: Cao, Yang, et al.
Veröffentlicht: (2024)
von: Cao, Yang, et al.
Veröffentlicht: (2024)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024)
Circuit Complexity Bounds for Visual Autoregressive Model
von: Ke, Yekun, et al.
Veröffentlicht: (2025)
von: Ke, Yekun, et al.
Veröffentlicht: (2025)
On Fine-Grained I/O Complexity of Attention Backward Passes
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024)
Unraveling the Smoothness Properties of Diffusion Models: A Gaussian Mixture Perspective
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
Modern Hopfield Networks Require Chain-of-Thought to Solve $\mathsf{NC}^1$-Hard Problems
von: Cao, Yang, et al.
Veröffentlicht: (2024)
von: Cao, Yang, et al.
Veröffentlicht: (2024)
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
Circuit Complexity Bounds for RoPE-based Transformer Architecture
von: Chen, Bo, et al.
Veröffentlicht: (2024)
von: Chen, Bo, et al.
Veröffentlicht: (2024)
HSR-Enhanced Sparse Attention Acceleration
von: Chen, Bo, et al.
Veröffentlicht: (2024)
von: Chen, Bo, et al.
Veröffentlicht: (2024)
Differential Privacy Mechanisms in Neural Tangent Kernel Regression
von: Gu, Jiuxiang, et al.
Veröffentlicht: (2024)
von: Gu, Jiuxiang, et al.
Veröffentlicht: (2024)
HOFAR: High-Order Augmentation of Flow Autoregressive Transformers
von: Liang, Yingyu, et al.
Veröffentlicht: (2025)
von: Liang, Yingyu, et al.
Veröffentlicht: (2025)
Beyond Linear Approximations: A Novel Pruning Approach for Attention Matrix
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent
von: Chen, Bo, et al.
Veröffentlicht: (2025)
von: Chen, Bo, et al.
Veröffentlicht: (2025)
Dissecting Submission Limit in Desk-Rejections: A Mathematical Analysis of Fairness in AI Conference Policies
von: Cao, Yuefan, et al.
Veröffentlicht: (2025)
von: Cao, Yuefan, et al.
Veröffentlicht: (2025)
RichSpace: Enriching Text-to-Video Prompt Space via Text Embedding Interpolation
von: Cao, Yuefan, et al.
Veröffentlicht: (2025)
von: Cao, Yuefan, et al.
Veröffentlicht: (2025)
Bypassing the Exponential Dependency: Looped Transformers Efficiently Learn In-context by Multi-step Gradient Descent
von: Chen, Bo, et al.
Veröffentlicht: (2024)
von: Chen, Bo, et al.
Veröffentlicht: (2024)
N-ReLU: Zero-Mean Stochastic Extension of ReLU
von: Manik, Md Motaleb Hossen, et al.
Veröffentlicht: (2025)
von: Manik, Md Motaleb Hossen, et al.
Veröffentlicht: (2025)
The Resurrection of the ReLU
von: Horuz, Coşku Can, et al.
Veröffentlicht: (2025)
von: Horuz, Coşku Can, et al.
Veröffentlicht: (2025)
Complexity of Injectivity and Verification of ReLU Neural Networks
von: Froese, Vincent, et al.
Veröffentlicht: (2024)
von: Froese, Vincent, et al.
Veröffentlicht: (2024)
High-Order Matching for One-Step Shortcut Diffusion Models
von: Chen, Bo, et al.
Veröffentlicht: (2025)
von: Chen, Bo, et al.
Veröffentlicht: (2025)
Force Matching with Relativistic Constraints: A Physics-Inspired Approach to Stable and Efficient Generative Modeling
von: Cao, Yang, et al.
Veröffentlicht: (2025)
von: Cao, Yang, et al.
Veröffentlicht: (2025)
The Geometry of ReLU Networks through the ReLU Transition Graph
von: Dhayalkar, Sahil Rajesh
Veröffentlicht: (2025)
von: Dhayalkar, Sahil Rajesh
Veröffentlicht: (2025)
Exploring the Frontiers of Softmax: Provable Optimization, Applications in Diffusion Model, and Beyond
von: Cao, Yang, et al.
Veröffentlicht: (2024)
von: Cao, Yang, et al.
Veröffentlicht: (2024)
Provable Differentially Private Computation of the Cross-Attention Mechanism
von: Ke, Yekun, et al.
Veröffentlicht: (2024)
von: Ke, Yekun, et al.
Veröffentlicht: (2024)
The Computational Complexity of Counting Linear Regions in ReLU Neural Networks
von: Stargalla, Moritz, et al.
Veröffentlicht: (2025)
von: Stargalla, Moritz, et al.
Veröffentlicht: (2025)
Training Tensor Attention Efficiently: From Cubic to Almost Linear Time
von: Cao, Yang, et al.
Veröffentlicht: (2024)
von: Cao, Yang, et al.
Veröffentlicht: (2024)
Towards Infinite-Long Prefix in Transformer
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
ReLU Neural Networks of Polynomial Size for Exact Maximum Flow Computation
von: Hertrich, Christoph, et al.
Veröffentlicht: (2021)
von: Hertrich, Christoph, et al.
Veröffentlicht: (2021)
Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformers
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
Training Cross-Morphology Embodied AI Agents: From Practical Challenges to Theoretical Foundations
von: Liu, Shaoshan, et al.
Veröffentlicht: (2025)
von: Liu, Shaoshan, et al.
Veröffentlicht: (2025)
A Tighter Complexity Analysis of SparseGPT
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024)
Do Large Language Models Have Compositional Ability? An Investigation into Limitations and Scalability
von: Xu, Zhuoyan, et al.
Veröffentlicht: (2024)
von: Xu, Zhuoyan, et al.
Veröffentlicht: (2024)
Is ReLU Adversarially Robust?
von: Sooksatra, Korn, et al.
Veröffentlicht: (2024)
von: Sooksatra, Korn, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis
von: Ke, Yekun, et al.
Veröffentlicht: (2025) -
Multi-Layer Transformers Gradient Can be Approximated in Almost Linear Time
von: Liang, Yingyu, et al.
Veröffentlicht: (2024) -
On Computational Limits of FlowAR Models: Expressivity and Efficiency
von: Cao, Yang, et al.
Veröffentlicht: (2025) -
Neural Algorithmic Reasoning for Hypergraphs with Looped Transformers
von: Huang, Zekai, et al.
Veröffentlicht: (2025) -
When Can We Solve the Weighted Low Rank Approximation Problem in Truly Subquadratic Time?
von: Li, Chenyang, et al.
Veröffentlicht: (2025)