Understanding the Expressive Power and Mechanisms of Transformer for Sequence Modeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Mingze, E, Weinan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Expressive Power of Mixture-of-Experts for Structured Complex Tasks
von: Wang, Mingze, et al.
Veröffentlicht: (2025)
von: Wang, Mingze, et al.
Veröffentlicht: (2025)
How Transformers Get Rich: Approximation and Dynamics Analysis
von: Wang, Mingze, et al.
Veröffentlicht: (2024)
von: Wang, Mingze, et al.
Veröffentlicht: (2024)
GradPower: Powering Gradients for Faster Language Model Pre-Training
von: Wang, Jinbo, et al.
Veröffentlicht: (2025)
von: Wang, Jinbo, et al.
Veröffentlicht: (2025)
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
von: Wang, Jinbo, et al.
Veröffentlicht: (2025)
von: Wang, Jinbo, et al.
Veröffentlicht: (2025)
On the Expressive Power of Floating-Point Transformers
von: Park, Sejun, et al.
Veröffentlicht: (2026)
von: Park, Sejun, et al.
Veröffentlicht: (2026)
On the Expressive Power of Contextual Relations in Transformers
von: Fraiman, Demián
Veröffentlicht: (2026)
von: Fraiman, Demián
Veröffentlicht: (2026)
Expressivity-Efficiency Tradeoffs for Hybrid Sequence Models
von: Cooper, John, et al.
Veröffentlicht: (2026)
von: Cooper, John, et al.
Veröffentlicht: (2026)
Transformers are Expressive, But Are They Expressive Enough for Regression?
von: Nath, Swaroop, et al.
Veröffentlicht: (2024)
von: Nath, Swaroop, et al.
Veröffentlicht: (2024)
More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations
von: Wang, Mingze, et al.
Veröffentlicht: (2026)
von: Wang, Mingze, et al.
Veröffentlicht: (2026)
Exact Expressive Power of Transformers with Padding
von: Merrill, William, et al.
Veröffentlicht: (2025)
von: Merrill, William, et al.
Veröffentlicht: (2025)
Understanding and Enhancing Mask-Based Pretraining towards Universal Representations
von: Dong, Mingze, et al.
Veröffentlicht: (2025)
von: Dong, Mingze, et al.
Veröffentlicht: (2025)
The Expressive Power of Transformers with Chain of Thought
von: Merrill, William, et al.
Veröffentlicht: (2023)
von: Merrill, William, et al.
Veröffentlicht: (2023)
On Expressive Power of Looped Transformers: Theoretical Analysis and Enhancement via Timestep Encoding
von: Xu, Kevin, et al.
Veröffentlicht: (2024)
von: Xu, Kevin, et al.
Veröffentlicht: (2024)
On The Expressive Power of GNN Derivatives
von: Eitan, Yam, et al.
Veröffentlicht: (2025)
von: Eitan, Yam, et al.
Veröffentlicht: (2025)
Lyra: An Efficient and Expressive Subquadratic Architecture for Modeling Biological Sequences
von: Ramesh, Krithik, et al.
Veröffentlicht: (2025)
von: Ramesh, Krithik, et al.
Veröffentlicht: (2025)
Structured Linear CDEs: Maximally Expressive and Parallel-in-Time Sequence Models
von: Walker, Benjamin, et al.
Veröffentlicht: (2025)
von: Walker, Benjamin, et al.
Veröffentlicht: (2025)
Towards Understanding the Expressive Power of GNNs with Global Readout
von: Funk, Maurice, et al.
Veröffentlicht: (2026)
von: Funk, Maurice, et al.
Veröffentlicht: (2026)
On the Theoretical Expressive Power and the Design Space of Higher-Order Graph Transformers
von: Zhou, Cai, et al.
Veröffentlicht: (2024)
von: Zhou, Cai, et al.
Veröffentlicht: (2024)
On the Expressive Power of Transformers for Maxout Networks and Continuous Piecewise Linear Functions
von: Gu, Linyan, et al.
Veröffentlicht: (2026)
von: Gu, Linyan, et al.
Veröffentlicht: (2026)
Expressive Power of Temporal Message Passing
von: Wałęga, Przemysław Andrzej, et al.
Veröffentlicht: (2024)
von: Wałęga, Przemysław Andrzej, et al.
Veröffentlicht: (2024)
Neural Attention: A Novel Mechanism for Enhanced Expressive Power in Transformer Models
von: DiGiugno, Andrew, et al.
Veröffentlicht: (2025)
von: DiGiugno, Andrew, et al.
Veröffentlicht: (2025)
Rethinking the Expressive Power of GNNs via Graph Biconnectivity
von: Zhang, Bohang, et al.
Veröffentlicht: (2023)
von: Zhang, Bohang, et al.
Veröffentlicht: (2023)
Expanding Expressivity in Transformer Models with MöbiusAttention
von: Halacheva, Anna-Maria, et al.
Veröffentlicht: (2024)
von: Halacheva, Anna-Maria, et al.
Veröffentlicht: (2024)
GNNs Meet Sequence Models Along the Shortest-Path: an Expressive Method for Link Prediction
von: Ferrini, Francesco, et al.
Veröffentlicht: (2025)
von: Ferrini, Francesco, et al.
Veröffentlicht: (2025)
On the Expressive Power of GNNs to Solve Linear SDPs
von: Qian, Chendi, et al.
Veröffentlicht: (2026)
von: Qian, Chendi, et al.
Veröffentlicht: (2026)
The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought
von: Brösamle, Moritz, et al.
Veröffentlicht: (2026)
von: Brösamle, Moritz, et al.
Veröffentlicht: (2026)
On the Expressive Power of Subgraph Graph Neural Networks for Graphs with Bounded Cycles
von: Chen, Ziang, et al.
Veröffentlicht: (2025)
von: Chen, Ziang, et al.
Veröffentlicht: (2025)
k-Maximum Inner Product Attention for Graph Transformers and the Expressive Power of GraphGPS
von: De Schouwer, Jonas, et al.
Veröffentlicht: (2026)
von: De Schouwer, Jonas, et al.
Veröffentlicht: (2026)
Understanding Expressivity of GNN in Rule Learning
von: Qiu, Haiquan, et al.
Veröffentlicht: (2023)
von: Qiu, Haiquan, et al.
Veröffentlicht: (2023)
Expressivity of Transformers: A Tropical Geometry Perspective
von: Su, Ye, et al.
Veröffentlicht: (2026)
von: Su, Ye, et al.
Veröffentlicht: (2026)
On the Expressive Power of Graph Neural Networks
von: Nalwade, Ashwin, et al.
Veröffentlicht: (2024)
von: Nalwade, Ashwin, et al.
Veröffentlicht: (2024)
On the Expressive Power of Sparse Geometric MPNNs
von: Sverdlov, Yonatan, et al.
Veröffentlicht: (2024)
von: Sverdlov, Yonatan, et al.
Veröffentlicht: (2024)
On the Expressive Power of GNNs for Boolean Satisfiability
von: Peltonen, Saku, et al.
Veröffentlicht: (2026)
von: Peltonen, Saku, et al.
Veröffentlicht: (2026)
On the Expressive Power of Permutation-Equivariant Weight-Space Networks
von: Dayan, Adir, et al.
Veröffentlicht: (2026)
von: Dayan, Adir, et al.
Veröffentlicht: (2026)
Introduction to Sequence Modeling with Transformers
von: Kämäräinen, Joni-Kristian
Veröffentlicht: (2025)
von: Kämäräinen, Joni-Kristian
Veröffentlicht: (2025)
Towards Dynamic Graph Neural Networks with Provably High-Order Expressive Power
von: Wang, Zhe, et al.
Veröffentlicht: (2024)
von: Wang, Zhe, et al.
Veröffentlicht: (2024)
A Little Depth Goes a Long Way: The Expressive Power of Log-Depth Transformers
von: Merrill, William, et al.
Veröffentlicht: (2025)
von: Merrill, William, et al.
Veröffentlicht: (2025)
Maximising Quantum-Computing Expressive Power through Randomised Circuits
von: Yang, Yingli, et al.
Veröffentlicht: (2023)
von: Yang, Yingli, et al.
Veröffentlicht: (2023)
On the Expressive Power and Limitations of Multi-Layer SSMs
von: Zubić, Nikola, et al.
Veröffentlicht: (2026)
von: Zubić, Nikola, et al.
Veröffentlicht: (2026)
On the Expressive Power of Tree-Structured Probabilistic Circuits
von: Yin, Lang, et al.
Veröffentlicht: (2024)
von: Yin, Lang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
On the Expressive Power of Mixture-of-Experts for Structured Complex Tasks
von: Wang, Mingze, et al.
Veröffentlicht: (2025) -
How Transformers Get Rich: Approximation and Dynamics Analysis
von: Wang, Mingze, et al.
Veröffentlicht: (2024) -
GradPower: Powering Gradients for Faster Language Model Pre-Training
von: Wang, Jinbo, et al.
Veröffentlicht: (2025) -
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
von: Wang, Jinbo, et al.
Veröffentlicht: (2025) -
On the Expressive Power of Floating-Point Transformers
von: Park, Sejun, et al.
Veröffentlicht: (2026)