Extra Global Attention Designation Using Keyword Detection in Sparse Transformer Architectures
Fuente:
arXiv
Salvato in:
| Autori principali: | Lucas, Evan, Kangas, Dylan, Havens, Timothy C |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Optimizing Native Sparse Attention with Latent Attention and Local Global Alternating Strategies
di: Hu, Yuxuan, et al.
Pubblicazione: (2025)
di: Hu, Yuxuan, et al.
Pubblicazione: (2025)
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
di: Nawrot, Piotr, et al.
Pubblicazione: (2025)
di: Nawrot, Piotr, et al.
Pubblicazione: (2025)
Sparse Growing Transformer: Training-Time Sparse Depth Allocation via Progressive Attention Looping
di: Chen, Yao, et al.
Pubblicazione: (2026)
di: Chen, Yao, et al.
Pubblicazione: (2026)
Data-Informed Global Sparseness in Attention Mechanisms for Deep Neural Networks
di: Rugina, Ileana, et al.
Pubblicazione: (2020)
di: Rugina, Ileana, et al.
Pubblicazione: (2020)
SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space
di: Shen, Zhenyi, et al.
Pubblicazione: (2025)
di: Shen, Zhenyi, et al.
Pubblicazione: (2025)
The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
di: Fichtl, Alexander M., et al.
Pubblicazione: (2025)
di: Fichtl, Alexander M., et al.
Pubblicazione: (2025)
HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
di: Gao, Yizhao, et al.
Pubblicazione: (2026)
di: Gao, Yizhao, et al.
Pubblicazione: (2026)
Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models
di: Liang, Weixin, et al.
Pubblicazione: (2024)
di: Liang, Weixin, et al.
Pubblicazione: (2024)
Rectified Sparse Attention
di: Sun, Yutao, et al.
Pubblicazione: (2025)
di: Sun, Yutao, et al.
Pubblicazione: (2025)
SparseD: Sparse Attention for Diffusion Language Models
di: Wang, Zeqing, et al.
Pubblicazione: (2025)
di: Wang, Zeqing, et al.
Pubblicazione: (2025)
$π$-Attention: Periodic Sparse Transformers for Efficient Long-Context Modeling
di: Liu, Dong, et al.
Pubblicazione: (2025)
di: Liu, Dong, et al.
Pubblicazione: (2025)
From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency
di: Wen, Kaiyue, et al.
Pubblicazione: (2024)
di: Wen, Kaiyue, et al.
Pubblicazione: (2024)
Chaos with Keywords: Exposing Large Language Models Sycophantic Hallucination to Misleading Keywords and Evaluating Defense Strategies
di: RRV, Aswin, et al.
Pubblicazione: (2024)
di: RRV, Aswin, et al.
Pubblicazione: (2024)
Keyword-Centric Prompting for One-Shot Event Detection with Self-Generated Rationale Enhancements
di: Li, Ziheng, et al.
Pubblicazione: (2025)
di: Li, Ziheng, et al.
Pubblicazione: (2025)
Tag&Tab: Pretraining Data Detection in Large Language Models Using Keyword-Based Membership Inference Attack
di: Antebi, Sagiv, et al.
Pubblicazione: (2025)
di: Antebi, Sagiv, et al.
Pubblicazione: (2025)
SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs
di: Gao, Yizhao, et al.
Pubblicazione: (2024)
di: Gao, Yizhao, et al.
Pubblicazione: (2024)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
di: Lou, Chao, et al.
Pubblicazione: (2024)
di: Lou, Chao, et al.
Pubblicazione: (2024)
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
di: Wang, Junxuan, et al.
Pubblicazione: (2025)
di: Wang, Junxuan, et al.
Pubblicazione: (2025)
FASA: Frequency-aware Sparse Attention
di: Wang, Yifei, et al.
Pubblicazione: (2026)
di: Wang, Yifei, et al.
Pubblicazione: (2026)
A Transformer with Stack Attention
di: Li, Jiaoda, et al.
Pubblicazione: (2024)
di: Li, Jiaoda, et al.
Pubblicazione: (2024)
BIG-Bench Extra Hard
di: Kazemi, Mehran, et al.
Pubblicazione: (2025)
di: Kazemi, Mehran, et al.
Pubblicazione: (2025)
Block-Wise Differentiable Sinkhorn Attention: Tail-Refinement Gradients with a Gap-Aware Dustbin Bridge
di: Forde, Dylan
Pubblicazione: (2026)
di: Forde, Dylan
Pubblicazione: (2026)
Beneath the Surface: The Role of Underwater Image Enhancement in Object Detection
di: Awad, Ali, et al.
Pubblicazione: (2024)
di: Awad, Ali, et al.
Pubblicazione: (2024)
Improving Performance of Automatic Keyword Extraction (AKE) Methods Using PoS-Tagging and Enhanced Semantic-Awareness
di: Altuncu, Enes, et al.
Pubblicazione: (2022)
di: Altuncu, Enes, et al.
Pubblicazione: (2022)
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
di: Piękos, Piotr, et al.
Pubblicazione: (2025)
di: Piękos, Piotr, et al.
Pubblicazione: (2025)
SEA: Sparse Linear Attention with Estimated Attention Mask
di: Lee, Heejun, et al.
Pubblicazione: (2023)
di: Lee, Heejun, et al.
Pubblicazione: (2023)
Revisiting Funnel Transformers for Modern LLM Architectures with Comprehensive Ablations in Training and Inference Configurations
di: Choi, DongHyun, et al.
Pubblicazione: (2025)
di: Choi, DongHyun, et al.
Pubblicazione: (2025)
ROUGE-K: Do Your Summaries Have Keywords?
di: Takeshita, Sotaro, et al.
Pubblicazione: (2024)
di: Takeshita, Sotaro, et al.
Pubblicazione: (2024)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
di: Yuan, Jingyang, et al.
Pubblicazione: (2025)
di: Yuan, Jingyang, et al.
Pubblicazione: (2025)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
di: Deng, Yichuan, et al.
Pubblicazione: (2024)
di: Deng, Yichuan, et al.
Pubblicazione: (2024)
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
di: Wang, Hanrui, et al.
Pubblicazione: (2020)
di: Wang, Hanrui, et al.
Pubblicazione: (2020)
Understanding In-Context Learning Beyond Transformers: An Investigation of State Space and Hybrid Architectures
di: Wang, Shenran, et al.
Pubblicazione: (2025)
di: Wang, Shenran, et al.
Pubblicazione: (2025)
SEKE: Specialised Experts for Keyword Extraction
di: Martinc, Matej, et al.
Pubblicazione: (2024)
di: Martinc, Matej, et al.
Pubblicazione: (2024)
Efficient Sparse Attention needs Adaptive Token Release
di: Zhang, Chaoran, et al.
Pubblicazione: (2024)
di: Zhang, Chaoran, et al.
Pubblicazione: (2024)
Lag-Relative Sparse Attention In Long Context Training
di: Liang, Manlai, et al.
Pubblicazione: (2025)
di: Liang, Manlai, et al.
Pubblicazione: (2025)
Adapting General-Purpose Embedding Models to Private Datasets Using Keyword-based Retrieval
di: Wei, Yubai, et al.
Pubblicazione: (2025)
di: Wei, Yubai, et al.
Pubblicazione: (2025)
Block Sparse Flash Attention
di: Ohayon, Daniel, et al.
Pubblicazione: (2025)
di: Ohayon, Daniel, et al.
Pubblicazione: (2025)
Long-Context Generalization with Sparse Attention
di: Vasylenko, Pavlo, et al.
Pubblicazione: (2025)
di: Vasylenko, Pavlo, et al.
Pubblicazione: (2025)
For those who don't know (how) to ask: Building a dataset of technology questions for digital newcomers
di: Lucas, Evan, et al.
Pubblicazione: (2024)
di: Lucas, Evan, et al.
Pubblicazione: (2024)
Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
di: Yan, Siyuan, et al.
Pubblicazione: (2025)
di: Yan, Siyuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Optimizing Native Sparse Attention with Latent Attention and Local Global Alternating Strategies
di: Hu, Yuxuan, et al.
Pubblicazione: (2025) -
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
di: Nawrot, Piotr, et al.
Pubblicazione: (2025) -
Sparse Growing Transformer: Training-Time Sparse Depth Allocation via Progressive Attention Looping
di: Chen, Yao, et al.
Pubblicazione: (2026) -
Data-Informed Global Sparseness in Attention Mechanisms for Deep Neural Networks
di: Rugina, Ileana, et al.
Pubblicazione: (2020) -
SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space
di: Shen, Zhenyi, et al.
Pubblicazione: (2025)