Fast attention mechanisms: a tale of parallelism
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Jingwen, Yu, Hantao, Sanford, Clayton, Andoni, Alexandr, Hsu, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fixed Universal Transformers
von: Liu, Jingwen, et al.
Veröffentlicht: (2026)
von: Liu, Jingwen, et al.
Veröffentlicht: (2026)
Transformers, parallel computation, and logarithmic depth
von: Sanford, Clayton, et al.
Veröffentlicht: (2024)
von: Sanford, Clayton, et al.
Veröffentlicht: (2024)
One-layer transformers fail to solve the induction heads task
von: Sanford, Clayton, et al.
Veröffentlicht: (2024)
von: Sanford, Clayton, et al.
Veröffentlicht: (2024)
A multi-source data power load forecasting method using attention mechanism-based parallel cnn-gru
von: Min, Chao, et al.
Veröffentlicht: (2024)
von: Min, Chao, et al.
Veröffentlicht: (2024)
When Do Transformers Outperform Feedforward and Recurrent Networks? A Statistical Perspective
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2025)
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2025)
Group-realizable multi-group learning by minimizing empirical risk
von: Ardeshir, Navid, et al.
Veröffentlicht: (2026)
von: Ardeshir, Navid, et al.
Veröffentlicht: (2026)
Group-wise oracle-efficient algorithms for online multi-group learning
von: Deng, Samuel, et al.
Veröffentlicht: (2024)
von: Deng, Samuel, et al.
Veröffentlicht: (2024)
Easy attention: A simple attention mechanism for temporal predictions with transformers
von: Sanchis-Agudo, Marcial, et al.
Veröffentlicht: (2023)
von: Sanchis-Agudo, Marcial, et al.
Veröffentlicht: (2023)
Approximation of relation functions and attention mechanisms
von: Altabaa, Awni, et al.
Veröffentlicht: (2024)
von: Altabaa, Awni, et al.
Veröffentlicht: (2024)
Flow Straight and Fast in Hilbert Space: Functional Rectified Flow
von: Zhang, Jianxin, et al.
Veröffentlicht: (2025)
von: Zhang, Jianxin, et al.
Veröffentlicht: (2025)
FMamba: Mamba based on Fast-attention for Multivariate Time-series Forecasting
von: Ma, Shusen, et al.
Veröffentlicht: (2024)
von: Ma, Shusen, et al.
Veröffentlicht: (2024)
Lost in Tokenization: Fundamental Trade-offs in Graph Tokenization for Transformers
von: Bechler-Speicher, Maya, et al.
Veröffentlicht: (2026)
von: Bechler-Speicher, Maya, et al.
Veröffentlicht: (2026)
Reservoir observer enhanced with residual calibration and attention mechanism
von: Liu, Yichen, et al.
Veröffentlicht: (2026)
von: Liu, Yichen, et al.
Veröffentlicht: (2026)
A foundation model with multi-variate parallel attention to generate neuronal activity
von: Carzaniga, Francesco, et al.
Veröffentlicht: (2025)
von: Carzaniga, Francesco, et al.
Veröffentlicht: (2025)
Two Heads Are Better than One: Simulating Large Transformers with Small Ones
von: Yu, Hantao, et al.
Veröffentlicht: (2025)
von: Yu, Hantao, et al.
Veröffentlicht: (2025)
Relational inductive biases on attention mechanisms
von: Mijangos, Víctor, et al.
Veröffentlicht: (2025)
von: Mijangos, Víctor, et al.
Veröffentlicht: (2025)
Next-Token Prediction and Regret Minimization
von: Mohri, Mehryar, et al.
Veröffentlicht: (2026)
von: Mohri, Mehryar, et al.
Veröffentlicht: (2026)
Best of Both Worlds: Advantages of Hybrid Graph Sequence Models
von: Behrouz, Ali, et al.
Veröffentlicht: (2024)
von: Behrouz, Ali, et al.
Veröffentlicht: (2024)
Parent-Guided Semantic Reward Model (PGSRM): Embedding-Based Reward Functions for Reinforcement Learning of Transformer Language Models
von: Plashchinsky, Alexandr
Veröffentlicht: (2025)
von: Plashchinsky, Alexandr
Veröffentlicht: (2025)
Fundamental Limitations on Subquadratic Alternatives to Transformers
von: Alman, Josh, et al.
Veröffentlicht: (2024)
von: Alman, Josh, et al.
Veröffentlicht: (2024)
Redundant feature screening method for human activity recognition based on attention purification mechanism
von: Li, Xiaoyang, et al.
Veröffentlicht: (2025)
von: Li, Xiaoyang, et al.
Veröffentlicht: (2025)
Implicit Bias and Fast Convergence Rates for Self-attention
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2024)
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2024)
A First Guess is Rarely the Final Answer: Learning to Search in the Traveling Salesperson Problem
von: Garmendia, Andoni Irazusta
Veröffentlicht: (2026)
von: Garmendia, Andoni Irazusta
Veröffentlicht: (2026)
Tucker Attention: A generalization of approximate attention mechanisms
von: Klein, Timon, et al.
Veröffentlicht: (2026)
von: Klein, Timon, et al.
Veröffentlicht: (2026)
Fast parallel sampling under isoperimetry
von: Anari, Nima, et al.
Veröffentlicht: (2024)
von: Anari, Nima, et al.
Veröffentlicht: (2024)
Inexact calculus of variations on the hyperspherical tangent bundle and its connections to the attention mechanism
von: Gracyk, Andrew
Veröffentlicht: (2025)
von: Gracyk, Andrew
Veröffentlicht: (2025)
Hierarchical Motion Captioning Utilizing External Text Data Source
von: Leite, Clayton, et al.
Veröffentlicht: (2025)
von: Leite, Clayton, et al.
Veröffentlicht: (2025)
Enhancing Motion Variation in Text-to-Motion Models via Pose and Video Conditioned Editing
von: Leite, Clayton, et al.
Veröffentlicht: (2024)
von: Leite, Clayton, et al.
Veröffentlicht: (2024)
Statistical-Computational Trade-offs for Density Estimation
von: Aamand, Anders, et al.
Veröffentlicht: (2024)
von: Aamand, Anders, et al.
Veröffentlicht: (2024)
Two tales for a geometric Jensen--Shannon divergence
von: Nielsen, Frank
Veröffentlicht: (2025)
von: Nielsen, Frank
Veröffentlicht: (2025)
Scaling Federated Linear Contextual Bandits via Sketching
von: Yang, Hantao, et al.
Veröffentlicht: (2026)
von: Yang, Hantao, et al.
Veröffentlicht: (2026)
Understanding Transformer Reasoning Capabilities via Graph Algorithms
von: Sanford, Clayton, et al.
Veröffentlicht: (2024)
von: Sanford, Clayton, et al.
Veröffentlicht: (2024)
Depth-Width tradeoffs in Algorithmic Reasoning of Graph Tasks with Transformers
von: Yehudai, Gilad, et al.
Veröffentlicht: (2025)
von: Yehudai, Gilad, et al.
Veröffentlicht: (2025)
Sink vs. diagonal patterns as mechanisms for attention switch and oversmoothing prevention
von: Súkeník, Peter, et al.
Veröffentlicht: (2026)
von: Súkeník, Peter, et al.
Veröffentlicht: (2026)
Invertible Memory Flow Networks
von: Zerihun, Liyu, et al.
Veröffentlicht: (2026)
von: Zerihun, Liyu, et al.
Veröffentlicht: (2026)
Lower bounds for one-layer transformers that compute parity
von: Hsu, Daniel
Veröffentlicht: (2026)
von: Hsu, Daniel
Veröffentlicht: (2026)
Locality Preserving Markovian Transition for Instance Retrieval
von: Luo, Jifei, et al.
Veröffentlicht: (2025)
von: Luo, Jifei, et al.
Veröffentlicht: (2025)
On the Computational Hardness of Transformers
von: Saha, Barna, et al.
Veröffentlicht: (2026)
von: Saha, Barna, et al.
Veröffentlicht: (2026)
Mapping of attention mechanisms to a generalized Potts model
von: Rende, Riccardo, et al.
Veröffentlicht: (2023)
von: Rende, Riccardo, et al.
Veröffentlicht: (2023)
Self-attentive Transformer for Fast and Accurate Postprocessing of Temperature and Wind Speed Forecasts
von: Van Poecke, Aaron, et al.
Veröffentlicht: (2024)
von: Van Poecke, Aaron, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Fixed Universal Transformers
von: Liu, Jingwen, et al.
Veröffentlicht: (2026) -
Transformers, parallel computation, and logarithmic depth
von: Sanford, Clayton, et al.
Veröffentlicht: (2024) -
One-layer transformers fail to solve the induction heads task
von: Sanford, Clayton, et al.
Veröffentlicht: (2024) -
A multi-source data power load forecasting method using attention mechanism-based parallel cnn-gru
von: Min, Chao, et al.
Veröffentlicht: (2024) -
When Do Transformers Outperform Feedforward and Recurrent Networks? A Statistical Perspective
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2025)