GatedFWA: Linear Flash Windowed Attention with Gated Associative Memory
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Liu, Jiaxu, Bai, Yuhe, Yin, Xiangyu, Bouganis, Christos-Savvas |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ATHEENA: A Toolflow for Hardware Early-Exit Network Automation
par: Biggs, Benjamin, et autres
Publié: (2023)
par: Biggs, Benjamin, et autres
Publié: (2023)
fpgaHART: A toolflow for throughput-oriented acceleration of 3D CNNs for HAR onto FPGAs
par: Toupas, Petros, et autres
Publié: (2023)
par: Toupas, Petros, et autres
Publié: (2023)
FMM-X3D: FPGA-based modeling and mapping of X3D for Human Action Recognition
par: Toupas, Petros, et autres
Publié: (2023)
par: Toupas, Petros, et autres
Publié: (2023)
Gated-SwinRMT: Unifying Swin Windowed Attention with Retentive Manhattan Decay via Input-Dependent Gating
par: Maity, Dipan, et autres
Publié: (2026)
par: Maity, Dipan, et autres
Publié: (2026)
Towards Understanding Why Label Smoothing Degrades Selective Classification and How to Fix It
par: Xia, Guoxuan, et autres
Publié: (2024)
par: Xia, Guoxuan, et autres
Publié: (2024)
SMOF: Streaming Modern CNNs on FPGAs with Smart Off-Chip Eviction
par: Toupas, Petros, et autres
Publié: (2024)
par: Toupas, Petros, et autres
Publié: (2024)
HARFLOW3D: A Latency-Oriented 3D-CNN Accelerator Toolflow for HAR on FPGA Devices
par: Toupas, Petros, et autres
Publié: (2023)
par: Toupas, Petros, et autres
Publié: (2023)
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
par: Li, Yingcong, et autres
Publié: (2025)
par: Li, Yingcong, et autres
Publié: (2025)
Gated Linear Attention Transformers with Hardware-Efficient Training
par: Yang, Songlin, et autres
Publié: (2023)
par: Yang, Songlin, et autres
Publié: (2023)
Differential Gated Self-Attention
par: Lygizou, Elpiniki Maria, et autres
Publié: (2025)
par: Lygizou, Elpiniki Maria, et autres
Publié: (2025)
A3 : an Analytical Low-Rank Approximation Framework for Attention
par: Wong, Jeffrey T. H., et autres
Publié: (2025)
par: Wong, Jeffrey T. H., et autres
Publié: (2025)
Emergent Self-Attention from Astrocyte-Gated Associative Memory Dynamics
par: Vivet, Arnau, et autres
Publié: (2026)
par: Vivet, Arnau, et autres
Publié: (2026)
Learnability Window in Gated Recurrent Neural Networks
par: Livi, Lorenzo
Publié: (2025)
par: Livi, Lorenzo
Publié: (2025)
Masked Gated Linear Unit
par: Tajima, Yukito, et autres
Publié: (2025)
par: Tajima, Yukito, et autres
Publié: (2025)
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
par: De, Soham, et autres
Publié: (2024)
par: De, Soham, et autres
Publié: (2024)
Gated Graph Attention Networks with Learnable Temperature
par: Ma, Zhongtian, et autres
Publié: (2026)
par: Ma, Zhongtian, et autres
Publié: (2026)
VeriGate: Verifier-Gated Step-Level Supervision for GRPO
par: Agrawal, Aakriti, et autres
Publié: (2026)
par: Agrawal, Aakriti, et autres
Publié: (2026)
Gated Associative Memory: A Parallel O(N) Architecture for Efficient Sequence Modeling
par: Acharya, Rishiraj
Publié: (2025)
par: Acharya, Rishiraj
Publié: (2025)
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
par: Chen, Shimao, et autres
Publié: (2024)
par: Chen, Shimao, et autres
Publié: (2024)
On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation
par: Diep, Nghiem T., et autres
Publié: (2025)
par: Diep, Nghiem T., et autres
Publié: (2025)
Quadratic Gating Mixture of Experts: Statistical Insights into Self-Attention
par: Akbarian, Pedram, et autres
Publié: (2024)
par: Akbarian, Pedram, et autres
Publié: (2024)
Boosting House Price Estimations with Multi-Head Gated Attention
par: Sellam, Zakaria Abdellah, et autres
Publié: (2024)
par: Sellam, Zakaria Abdellah, et autres
Publié: (2024)
Gating Enables Curvature: A Geometric Expressivity Gap in Attention
par: Bathula, Satwik, et autres
Publié: (2026)
par: Bathula, Satwik, et autres
Publié: (2026)
xLSTM-PINN: Memory-Gated Spectral Remodeling for Physics-Informed Learning
par: Tao, Ze, et autres
Publié: (2025)
par: Tao, Ze, et autres
Publié: (2025)
Variational Linear Attention: Stable Associative Memory for Long-Context Transformers
par: Pandey, Vishal, et autres
Publié: (2026)
par: Pandey, Vishal, et autres
Publié: (2026)
Tiled Flash Linear Attention: More Efficient Linear RNN and xLSTM Kernels
par: Beck, Maximilian, et autres
Publié: (2025)
par: Beck, Maximilian, et autres
Publié: (2025)
FineGates: LLMs Finetuning with Compression using Stochastic Gates
par: Svirsky, Jonathan, et autres
Publié: (2024)
par: Svirsky, Jonathan, et autres
Publié: (2024)
FlashMask: Efficient and Rich Mask Extension of FlashAttention
par: Wang, Guoxia, et autres
Publié: (2024)
par: Wang, Guoxia, et autres
Publié: (2024)
A Semantic Partitioning Method for Large-Scale Training of Knowledge Graph Embeddings
par: Bai, Yuhe
Publié: (2025)
par: Bai, Yuhe
Publié: (2025)
Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation
par: Zimerman, Itamar, et autres
Publié: (2024)
par: Zimerman, Itamar, et autres
Publié: (2024)
InterroGate: Learning to Share, Specialize, and Prune Representations for Multi-task Learning
par: Bejnordi, Babak Ehteshami, et autres
Publié: (2024)
par: Bejnordi, Babak Ehteshami, et autres
Publié: (2024)
HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator
par: Yu, Zhewen, et autres
Publié: (2024)
par: Yu, Zhewen, et autres
Publié: (2024)
In-context KV-Cache Eviction for LLMs via Attention-Gate
par: Zeng, Zihao, et autres
Publié: (2024)
par: Zeng, Zihao, et autres
Publié: (2024)
Efficient Learning for Linear Properties of Bounded-Gate Quantum Circuits
par: Du, Yuxuan, et autres
Publié: (2024)
par: Du, Yuxuan, et autres
Publié: (2024)
MOMEMTO: Patch-based Memory Gate Model in Time Series Foundation Model
par: Yoon, Samuel, et autres
Publié: (2025)
par: Yoon, Samuel, et autres
Publié: (2025)
Inductive Power Grid Cascading Failure Analysis with GRU-Gated Graph Attention
par: Zhou, Tianxin, et autres
Publié: (2026)
par: Zhou, Tianxin, et autres
Publié: (2026)
CGCMA: Conditionally-Gated Cross-Modal Attention for Event-Conditioned Asynchronous Fusion
par: Guo, Yunxiang
Publié: (2026)
par: Guo, Yunxiang
Publié: (2026)
A Statistical Theory of Gated Attention through the Lens of Hierarchical Mixture of Experts
par: Nguyen, Viet, et autres
Publié: (2026)
par: Nguyen, Viet, et autres
Publié: (2026)
Energy-Gated Attention: Spectral Salience as an Inductive Bias for Transformer Attention
par: Zeris, Athanasios
Publié: (2026)
par: Zeris, Athanasios
Publié: (2026)
Deconstructing Recurrence, Attention, and Gating: Investigating the transferability of Transformers and Gated Recurrent Neural Networks in forecasting of dynamical systems
par: Heidenreich, Hunter S., et autres
Publié: (2024)
par: Heidenreich, Hunter S., et autres
Publié: (2024)
Documents similaires
-
ATHEENA: A Toolflow for Hardware Early-Exit Network Automation
par: Biggs, Benjamin, et autres
Publié: (2023) -
fpgaHART: A toolflow for throughput-oriented acceleration of 3D CNNs for HAR onto FPGAs
par: Toupas, Petros, et autres
Publié: (2023) -
FMM-X3D: FPGA-based modeling and mapping of X3D for Human Action Recognition
par: Toupas, Petros, et autres
Publié: (2023) -
Gated-SwinRMT: Unifying Swin Windowed Attention with Retentive Manhattan Decay via Input-Dependent Gating
par: Maity, Dipan, et autres
Publié: (2026) -
Towards Understanding Why Label Smoothing Degrades Selective Classification and How to Fix It
par: Xia, Guoxuan, et autres
Publié: (2024)