Stack Attention: Improving the Ability of Transformers to Model Hierarchical Patterns
Fuente:
arXiv
Saved in:
| Main Authors: | DuSell, Brian, Chiang, David |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bearing Syntactic Fruit with Stack-Augmented Neural Networks
by: DuSell, Brian, et al.
Published: (2025)
by: DuSell, Brian, et al.
Published: (2025)
PILA: A Historical-Linguistic Dataset of Proto-Italic and Latin
by: Bothwell, Stephen, et al.
Published: (2024)
by: Bothwell, Stephen, et al.
Published: (2024)
Algorithms for Weighted Pushdown Automata
by: Butoi, Alexandra, et al.
Published: (2022)
by: Butoi, Alexandra, et al.
Published: (2022)
Information Locality as an Inductive Bias for Neural Language Models
by: Someya, Taiga, et al.
Published: (2025)
by: Someya, Taiga, et al.
Published: (2025)
On the Proper Treatment of Tokenization in Psycholinguistics
by: Giulianelli, Mario, et al.
Published: (2024)
by: Giulianelli, Mario, et al.
Published: (2024)
Training Neural Networks as Recognizers of Formal Languages
by: Butoi, Alexandra, et al.
Published: (2024)
by: Butoi, Alexandra, et al.
Published: (2024)
The Foundations of Tokenization: Statistical and Computational Concerns
by: Gastaldi, Juan Luis, et al.
Published: (2024)
by: Gastaldi, Juan Luis, et al.
Published: (2024)
From Language Models over Tokens to Language Models over Characters
by: Vieira, Tim, et al.
Published: (2024)
by: Vieira, Tim, et al.
Published: (2024)
A Transformer with Stack Attention
by: Li, Jiaoda, et al.
Published: (2024)
by: Li, Jiaoda, et al.
Published: (2024)
Language Models over Canonical Byte-Pair Encodings
by: Vieira, Tim, et al.
Published: (2025)
by: Vieira, Tim, et al.
Published: (2025)
Improving Rare Word Translation With Dictionaries and Attention Masking
by: Sible, Kenneth J., et al.
Published: (2024)
by: Sible, Kenneth J., et al.
Published: (2024)
Counting Like Transformers: Compiling Temporal Counting Logic Into Softmax Transformers
by: Yang, Andy, et al.
Published: (2024)
by: Yang, Andy, et al.
Published: (2024)
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
by: Musat, Tiberiu
Published: (2024)
by: Musat, Tiberiu
Published: (2024)
Improving Conversational Abilities of Quantized Large Language Models via Direct Preference Alignment
by: Lee, Janghwan, et al.
Published: (2024)
by: Lee, Janghwan, et al.
Published: (2024)
Wait, but Tylenol is Acetaminophen... Investigating and Improving Language Models' Ability to Resist Requests for Misinformation
by: Chen, Shan, et al.
Published: (2024)
by: Chen, Shan, et al.
Published: (2024)
On the Ability of Transformers to Verify Plans
by: Sarrof, Yash, et al.
Published: (2026)
by: Sarrof, Yash, et al.
Published: (2026)
The Attentional White Bear Effect in Transformer Language Models
by: Ramnauth, Rebecca, et al.
Published: (2026)
by: Ramnauth, Rebecca, et al.
Published: (2026)
Simulating Hard Attention Using Soft Attention
by: Yang, Andy, et al.
Published: (2024)
by: Yang, Andy, et al.
Published: (2024)
Improve LLM-as-a-Judge Ability as a General Ability
by: Yu, Jiachen, et al.
Published: (2025)
by: Yu, Jiachen, et al.
Published: (2025)
Masked Hard-Attention Transformers Recognize Exactly the Star-Free Languages
by: Yang, Andy, et al.
Published: (2023)
by: Yang, Andy, et al.
Published: (2023)
Transformers in Uniform TC$^0$
by: Chiang, David
Published: (2024)
by: Chiang, David
Published: (2024)
Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
by: Leng, Jiaqi, et al.
Published: (2025)
by: Leng, Jiaqi, et al.
Published: (2025)
Knee-Deep in C-RASP: A Transformer Depth Hierarchy
by: Yang, Andy, et al.
Published: (2025)
by: Yang, Andy, et al.
Published: (2025)
Pre-trained Language Models Improve the Few-shot Prompt Ability of Decision Transformer
by: Yang, Yu, et al.
Published: (2024)
by: Yang, Yu, et al.
Published: (2024)
Selective Attention Improves Transformer
by: Leviathan, Yaniv, et al.
Published: (2024)
by: Leviathan, Yaniv, et al.
Published: (2024)
GLU Attention Improve Transformer
by: Wang, Zehao
Published: (2025)
by: Wang, Zehao
Published: (2025)
ComplexFormer: Disruptively Advancing Transformer Inference Ability via Head-Specific Complex Vector Attention
by: Shao, Jintian, et al.
Published: (2025)
by: Shao, Jintian, et al.
Published: (2025)
Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training
by: Du, Wenyu, et al.
Published: (2024)
by: Du, Wenyu, et al.
Published: (2024)
Code Pretraining Improves Entity Tracking Abilities of Language Models
by: Kim, Najoung, et al.
Published: (2024)
by: Kim, Najoung, et al.
Published: (2024)
We're Calling an Intervention: Exploring Fundamental Hurdles in Adapting Language Models to Nonstandard Text
by: Srivastava, Aarohi, et al.
Published: (2024)
by: Srivastava, Aarohi, et al.
Published: (2024)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
by: Mihaila, George
Published: (2026)
by: Mihaila, George
Published: (2026)
Improving LLM Abilities in Idiomatic Translation
by: Donthi, Sundesh, et al.
Published: (2024)
by: Donthi, Sundesh, et al.
Published: (2024)
Improving Transformers with Dynamically Composable Multi-Head Attention
by: Xiao, Da, et al.
Published: (2024)
by: Xiao, Da, et al.
Published: (2024)
Emergent Stack Representations in Modeling Counter Languages Using Transformers
by: Tiwari, Utkarsh, et al.
Published: (2025)
by: Tiwari, Utkarsh, et al.
Published: (2025)
Demystifying the Slash Pattern in Attention: The Role of RoPE
by: Cheng, Yuan, et al.
Published: (2026)
by: Cheng, Yuan, et al.
Published: (2026)
Entailed Opinion Matters: Improving the Fact-Checking Performance of Language Models by Relying on their Entailment Ability
by: Kumar, Gaurav, et al.
Published: (2025)
by: Kumar, Gaurav, et al.
Published: (2025)
Diversity of Thought Improves Reasoning Abilities of LLMs
by: Naik, Ranjita, et al.
Published: (2023)
by: Naik, Ranjita, et al.
Published: (2023)
Fine-Tuning and Evaluating Open-Source Large Language Models for the Army Domain
by: Ruiz, Daniel C., et al.
Published: (2024)
by: Ruiz, Daniel C., et al.
Published: (2024)
Nostra Domina at EvaLatin 2024: Improving Latin Polarity Detection through Data Augmentation
by: Bothwell, Stephen, et al.
Published: (2024)
by: Bothwell, Stephen, et al.
Published: (2024)
Extracting and Combining Abilities For Building Multi-lingual Ability-enhanced Large Language Models
by: Chen, Zhipeng, et al.
Published: (2024)
by: Chen, Zhipeng, et al.
Published: (2024)
Similar Items
-
Bearing Syntactic Fruit with Stack-Augmented Neural Networks
by: DuSell, Brian, et al.
Published: (2025) -
PILA: A Historical-Linguistic Dataset of Proto-Italic and Latin
by: Bothwell, Stephen, et al.
Published: (2024) -
Algorithms for Weighted Pushdown Automata
by: Butoi, Alexandra, et al.
Published: (2022) -
Information Locality as an Inductive Bias for Neural Language Models
by: Someya, Taiga, et al.
Published: (2025) -
On the Proper Treatment of Tokenization in Psycholinguistics
by: Giulianelli, Mario, et al.
Published: (2024)