LASER: Attention with Exponential Transformation
Fuente:
arXiv
Salvato in:
| Autori principali: | Duvvuri, Sai Surya, Dhillon, Inderjit S. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LUCID: Attention with Preconditioned Representations
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2026)
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2026)
LoRA Done RITE: Robust Invariant Transformation Equilibration for LoRA Optimization
di: Yen, Jui-Nan, et al.
Pubblicazione: (2024)
di: Yen, Jui-Nan, et al.
Pubblicazione: (2024)
Interleaved Head Attention
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2026)
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2026)
Two-stage LLM Fine-tuning with Less Specialization and More Generalization
di: Wang, Yihan, et al.
Pubblicazione: (2022)
di: Wang, Yihan, et al.
Pubblicazione: (2022)
MatFormer: Nested Transformer for Elastic Inference
di: Devvrit, et al.
Pubblicazione: (2023)
di: Devvrit, et al.
Pubblicazione: (2023)
The Art of Scaling Reinforcement Learning Compute for LLMs
di: Khatri, Devvrit, et al.
Pubblicazione: (2025)
di: Khatri, Devvrit, et al.
Pubblicazione: (2025)
PowerAttention: Exponentially Scaling of Receptive Fields for Effective Sparse Attention
di: Chen, Lida, et al.
Pubblicazione: (2025)
di: Chen, Lida, et al.
Pubblicazione: (2025)
Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
di: Bansal, Rachit, et al.
Pubblicazione: (2025)
di: Bansal, Rachit, et al.
Pubblicazione: (2025)
Large Language Models are Interpretable Learners
di: Wang, Ruochen, et al.
Pubblicazione: (2024)
di: Wang, Ruochen, et al.
Pubblicazione: (2024)
Generalized Probabilistic Attention Mechanism in Transformers
di: Heo, DongNyeong, et al.
Pubblicazione: (2024)
di: Heo, DongNyeong, et al.
Pubblicazione: (2024)
Rethinking Attention Output Projection: Structured Hadamard Transforms for Efficient Transformers
di: Aggarwal, Shubham, et al.
Pubblicazione: (2026)
di: Aggarwal, Shubham, et al.
Pubblicazione: (2026)
Improving Transformers with Dynamically Composable Multi-Head Attention
di: Xiao, Da, et al.
Pubblicazione: (2024)
di: Xiao, Da, et al.
Pubblicazione: (2024)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
di: Bozic, Vukasin, et al.
Pubblicazione: (2023)
di: Bozic, Vukasin, et al.
Pubblicazione: (2023)
Learning to Focus: Focal Attention for Selective and Scalable Transformers
di: Ram, Dhananjay, et al.
Pubblicazione: (2025)
di: Ram, Dhananjay, et al.
Pubblicazione: (2025)
Gated Linear Attention Transformers with Hardware-Efficient Training
di: Yang, Songlin, et al.
Pubblicazione: (2023)
di: Yang, Songlin, et al.
Pubblicazione: (2023)
Extracting Rule-based Descriptions of Attention Features in Transformers
di: Friedman, Dan, et al.
Pubblicazione: (2025)
di: Friedman, Dan, et al.
Pubblicazione: (2025)
Selective Attention: Enhancing Transformer through Principled Context Control
di: Zhang, Xuechen, et al.
Pubblicazione: (2024)
di: Zhang, Xuechen, et al.
Pubblicazione: (2024)
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
di: Musat, Tiberiu
Pubblicazione: (2024)
di: Musat, Tiberiu
Pubblicazione: (2024)
Transformer Based Linear Attention with Optimized GPU Kernel Implementation
di: Gerami, Armin, et al.
Pubblicazione: (2025)
di: Gerami, Armin, et al.
Pubblicazione: (2025)
SAP: Syntactic Attention Pruning for Transformer-based Language Models
di: Lee, Tzu-Yun, et al.
Pubblicazione: (2025)
di: Lee, Tzu-Yun, et al.
Pubblicazione: (2025)
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
di: Nawrot, Piotr, et al.
Pubblicazione: (2025)
di: Nawrot, Piotr, et al.
Pubblicazione: (2025)
Faster Transformer Decoding: N-gram Masked Self-Attention
di: Chelba, Ciprian, et al.
Pubblicazione: (2020)
di: Chelba, Ciprian, et al.
Pubblicazione: (2020)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
di: Brandon, William, et al.
Pubblicazione: (2024)
di: Brandon, William, et al.
Pubblicazione: (2024)
AMPLIFY:Attention-based Mixup for Performance Improvement and Label Smoothing in Transformer
di: Yang, Leixin, et al.
Pubblicazione: (2023)
di: Yang, Leixin, et al.
Pubblicazione: (2023)
PaTH Attention: Position Encoding via Accumulating Householder Transformations
di: Yang, Songlin, et al.
Pubblicazione: (2025)
di: Yang, Songlin, et al.
Pubblicazione: (2025)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
di: Mihaila, George
Pubblicazione: (2026)
di: Mihaila, George
Pubblicazione: (2026)
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
di: Dong, Yihe, et al.
Pubblicazione: (2025)
di: Dong, Yihe, et al.
Pubblicazione: (2025)
Attention Is Not All You Need: The Importance of Feedforward Networks in Transformer Models
di: Gerber, Isaac
Pubblicazione: (2025)
di: Gerber, Isaac
Pubblicazione: (2025)
Selective Attention Improves Transformer
di: Leviathan, Yaniv, et al.
Pubblicazione: (2024)
di: Leviathan, Yaniv, et al.
Pubblicazione: (2024)
LeetDecoding: A PyTorch Library for Exponentially Decaying Causal Linear Attention with CUDA Implementations
di: Wang, Jiaping, et al.
Pubblicazione: (2025)
di: Wang, Jiaping, et al.
Pubblicazione: (2025)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
di: Lou, Chao, et al.
Pubblicazione: (2024)
di: Lou, Chao, et al.
Pubblicazione: (2024)
Learnable Multi-Scale Wavelet Transformer: A Novel Alternative to Self-Attention
di: Kiruluta, Andrew, et al.
Pubblicazione: (2025)
di: Kiruluta, Andrew, et al.
Pubblicazione: (2025)
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
di: Wang, Junxuan, et al.
Pubblicazione: (2025)
di: Wang, Junxuan, et al.
Pubblicazione: (2025)
Geometric Median (GM) Matching for Robust Data Pruning
di: Acharya, Anish, et al.
Pubblicazione: (2024)
di: Acharya, Anish, et al.
Pubblicazione: (2024)
ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization
di: Patel, Nirmal, et al.
Pubblicazione: (2026)
di: Patel, Nirmal, et al.
Pubblicazione: (2026)
Zeroth-Order Sharpness-Aware Learning with Exponential Tilting
di: Gong, Xuchen, et al.
Pubblicazione: (2025)
di: Gong, Xuchen, et al.
Pubblicazione: (2025)
Energy-Gated Attention: Spectral Salience as an Inductive Bias for Transformer Attention
di: Zeris, Athanasios
Pubblicazione: (2026)
di: Zeris, Athanasios
Pubblicazione: (2026)
Beyond Self Attention: A Subquadratic Fourier Wavelet Transformer with Multi Modal Fusion
di: Kiruluta, Andrew, et al.
Pubblicazione: (2021)
di: Kiruluta, Andrew, et al.
Pubblicazione: (2021)
Causal Inference for Human-Language Model Collaboration
di: Zhang, Bohan, et al.
Pubblicazione: (2024)
di: Zhang, Bohan, et al.
Pubblicazione: (2024)
A2SF: Accumulative Attention Scoring with Forgetting Factor for Token Pruning in Transformer Decoder
di: Jo, Hyun-rae, et al.
Pubblicazione: (2024)
di: Jo, Hyun-rae, et al.
Pubblicazione: (2024)
Documenti analoghi
-
LUCID: Attention with Preconditioned Representations
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2026) -
LoRA Done RITE: Robust Invariant Transformation Equilibration for LoRA Optimization
di: Yen, Jui-Nan, et al.
Pubblicazione: (2024) -
Interleaved Head Attention
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2026) -
Two-stage LLM Fine-tuning with Less Specialization and More Generalization
di: Wang, Yihan, et al.
Pubblicazione: (2022) -
MatFormer: Nested Transformer for Elastic Inference
di: Devvrit, et al.
Pubblicazione: (2023)