Exclusive Self Attention
Fuente:
arXiv
Salvato in:
| Autore principale: | Zhai, Shuangfei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive Flows
di: Zhang, Ruixiang, et al.
Pubblicazione: (2025)
di: Zhang, Ruixiang, et al.
Pubblicazione: (2025)
AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
di: Song, Dinghong, et al.
Pubblicazione: (2025)
di: Song, Dinghong, et al.
Pubblicazione: (2025)
ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition
di: Ye, Lu, et al.
Pubblicazione: (2024)
di: Ye, Lu, et al.
Pubblicazione: (2024)
Faster Transformer Decoding: N-gram Masked Self-Attention
di: Chelba, Ciprian, et al.
Pubblicazione: (2020)
di: Chelba, Ciprian, et al.
Pubblicazione: (2020)
Aligning Attention with Human Rationales for Self-Explaining Hate Speech Detection
di: Eilertsen, Brage, et al.
Pubblicazione: (2025)
di: Eilertsen, Brage, et al.
Pubblicazione: (2025)
Learnable Multi-Scale Wavelet Transformer: A Novel Alternative to Self-Attention
di: Kiruluta, Andrew, et al.
Pubblicazione: (2025)
di: Kiruluta, Andrew, et al.
Pubblicazione: (2025)
Entropic-Time Inference: Self-Organizing Large Language Model Decoding Beyond Attention
di: Kiruluta, Andrew
Pubblicazione: (2026)
di: Kiruluta, Andrew
Pubblicazione: (2026)
Beyond Self Attention: A Subquadratic Fourier Wavelet Transformer with Multi Modal Fusion
di: Kiruluta, Andrew, et al.
Pubblicazione: (2021)
di: Kiruluta, Andrew, et al.
Pubblicazione: (2021)
Adaptive Two Sided Laplace Transforms: A Learnable, Interpretable, and Scalable Replacement for Self-Attention
di: Kiruluta, Andrew
Pubblicazione: (2025)
di: Kiruluta, Andrew
Pubblicazione: (2025)
Dynamic Rank Reinforcement Learning for Adaptive Low-Rank Multi-Head Self Attention in Large Language Models
di: Erden, Caner
Pubblicazione: (2025)
di: Erden, Caner
Pubblicazione: (2025)
Empirical Capacity Model for Self-Attention Neural Networks
di: Härmä, Aki, et al.
Pubblicazione: (2024)
di: Härmä, Aki, et al.
Pubblicazione: (2024)
Why Softmax Attention Outperforms Linear Attention
di: Deng, Yichuan, et al.
Pubblicazione: (2023)
di: Deng, Yichuan, et al.
Pubblicazione: (2023)
SEA: Sparse Linear Attention with Estimated Attention Mask
di: Lee, Heejun, et al.
Pubblicazione: (2023)
di: Lee, Heejun, et al.
Pubblicazione: (2023)
RecurFormer: Not All Transformer Heads Need Self-Attention
di: Yan, Ruiqing, et al.
Pubblicazione: (2024)
di: Yan, Ruiqing, et al.
Pubblicazione: (2024)
Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers
di: Wong, Liang Ze
Pubblicazione: (2025)
di: Wong, Liang Ze
Pubblicazione: (2025)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
di: Qiu, Quantong, et al.
Pubblicazione: (2026)
di: Qiu, Quantong, et al.
Pubblicazione: (2026)
Don't Pay Attention, PLANT It: Pretraining Attention via Learning-to-Rank
di: Roy, Debjyoti Saha, et al.
Pubblicazione: (2024)
di: Roy, Debjyoti Saha, et al.
Pubblicazione: (2024)
PowerAttention: Exponentially Scaling of Receptive Fields for Effective Sparse Attention
di: Chen, Lida, et al.
Pubblicazione: (2025)
di: Chen, Lida, et al.
Pubblicazione: (2025)
Learning to Attribute with Attention
di: Cohen-Wang, Benjamin, et al.
Pubblicazione: (2025)
di: Cohen-Wang, Benjamin, et al.
Pubblicazione: (2025)
Scale-invariant Attention
di: Anson, Ben, et al.
Pubblicazione: (2025)
di: Anson, Ben, et al.
Pubblicazione: (2025)
Breaking the Attention Bottleneck
di: Hilsenbek, Kalle
Pubblicazione: (2024)
di: Hilsenbek, Kalle
Pubblicazione: (2024)
Stochastic Attention: Connectome-Inspired Randomized Routing for Expressive Linear-Time Attention
di: Jin, Zehao, et al.
Pubblicazione: (2026)
di: Jin, Zehao, et al.
Pubblicazione: (2026)
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
di: Ildiz, M. Emrullah, et al.
Pubblicazione: (2024)
di: Ildiz, M. Emrullah, et al.
Pubblicazione: (2024)
Dynamics of Spontaneous Topic Changes in Next Token Prediction with Self-Attention
di: Jia, Mumin, et al.
Pubblicazione: (2025)
di: Jia, Mumin, et al.
Pubblicazione: (2025)
Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
di: Shyam, Vasudev, et al.
Pubblicazione: (2024)
di: Shyam, Vasudev, et al.
Pubblicazione: (2024)
Fast Multipole Attention: A Scalable Multilevel Attention Mechanism for Text and Images
di: Kang, Yanming, et al.
Pubblicazione: (2023)
di: Kang, Yanming, et al.
Pubblicazione: (2023)
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
di: Ma, Qingsen, et al.
Pubblicazione: (2026)
di: Ma, Qingsen, et al.
Pubblicazione: (2026)
Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models
di: Deng, Difan, et al.
Pubblicazione: (2026)
di: Deng, Difan, et al.
Pubblicazione: (2026)
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
di: He, Mutian, et al.
Pubblicazione: (2025)
di: He, Mutian, et al.
Pubblicazione: (2025)
Optimizing Mixture of Block Attention
di: Xiao, Guangxuan, et al.
Pubblicazione: (2025)
di: Xiao, Guangxuan, et al.
Pubblicazione: (2025)
Linear Attention Sequence Parallelism
di: Sun, Weigao, et al.
Pubblicazione: (2024)
di: Sun, Weigao, et al.
Pubblicazione: (2024)
Causal Attention with Lookahead Keys
di: Song, Zhuoqing, et al.
Pubblicazione: (2025)
di: Song, Zhuoqing, et al.
Pubblicazione: (2025)
LASER: Attention with Exponential Transformation
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2024)
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2024)
Multi-matrix Factorization Attention
di: Hu, Jingcheng, et al.
Pubblicazione: (2024)
di: Hu, Jingcheng, et al.
Pubblicazione: (2024)
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
di: Piękos, Piotr, et al.
Pubblicazione: (2025)
di: Piękos, Piotr, et al.
Pubblicazione: (2025)
Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction
di: Filipek, Adam
Pubblicazione: (2025)
di: Filipek, Adam
Pubblicazione: (2025)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
di: Bozic, Vukasin, et al.
Pubblicazione: (2023)
di: Bozic, Vukasin, et al.
Pubblicazione: (2023)
You Need Better Attention Priors
di: Litman, Elon, et al.
Pubblicazione: (2026)
di: Litman, Elon, et al.
Pubblicazione: (2026)
Coupled Query-Key Dynamics for Attention
di: Gahtan, Barak, et al.
Pubblicazione: (2026)
di: Gahtan, Barak, et al.
Pubblicazione: (2026)
Hardware-Efficient Attention for Fast Decoding
di: Zadouri, Ted, et al.
Pubblicazione: (2025)
di: Zadouri, Ted, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive Flows
di: Zhang, Ruixiang, et al.
Pubblicazione: (2025) -
AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
di: Song, Dinghong, et al.
Pubblicazione: (2025) -
ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition
di: Ye, Lu, et al.
Pubblicazione: (2024) -
Faster Transformer Decoding: N-gram Masked Self-Attention
di: Chelba, Ciprian, et al.
Pubblicazione: (2020) -
Aligning Attention with Human Rationales for Self-Explaining Hate Speech Detection
di: Eilertsen, Brage, et al.
Pubblicazione: (2025)