Extracting Rule-based Descriptions of Attention Features in Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Friedman, Dan, Bhaskar, Adithya, Wettig, Alexander, Chen, Danqi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Finding Transformer Circuits with Edge Pruning
by: Bhaskar, Adithya, et al.
Published: (2024)
by: Bhaskar, Adithya, et al.
Published: (2024)
The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models
by: Bhaskar, Adithya, et al.
Published: (2024)
by: Bhaskar, Adithya, et al.
Published: (2024)
Representing Rule-based Chatbots with Transformers
by: Friedman, Dan, et al.
Published: (2024)
by: Friedman, Dan, et al.
Published: (2024)
QuRating: Selecting High-Quality Data for Training Language Models
by: Wettig, Alexander, et al.
Published: (2024)
by: Wettig, Alexander, et al.
Published: (2024)
How to Train Long-Context Language Models (Effectively)
by: Gao, Tianyu, et al.
Published: (2024)
by: Gao, Tianyu, et al.
Published: (2024)
Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?
by: Bhaskar, Adithya, et al.
Published: (2025)
by: Bhaskar, Adithya, et al.
Published: (2025)
Improving Language Understanding from Screenshots
by: Gao, Tianyu, et al.
Published: (2024)
by: Gao, Tianyu, et al.
Published: (2024)
Continual Memorization of Factoids in Language Models
by: Chen, Howard, et al.
Published: (2024)
by: Chen, Howard, et al.
Published: (2024)
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
by: Razin, Noam, et al.
Published: (2024)
by: Razin, Noam, et al.
Published: (2024)
Interpretability Illusions in the Generalization of Simplified Models
by: Friedman, Dan, et al.
Published: (2023)
by: Friedman, Dan, et al.
Published: (2023)
Language Models that Think, Chat Better
by: Bhaskar, Adithya, et al.
Published: (2025)
by: Bhaskar, Adithya, et al.
Published: (2025)
Learning to Extract Structured Entities Using Language Models
by: Wu, Haolun, et al.
Published: (2024)
by: Wu, Haolun, et al.
Published: (2024)
SimPO: Simple Preference Optimization with a Reference-Free Reward
by: Meng, Yu, et al.
Published: (2024)
by: Meng, Yu, et al.
Published: (2024)
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
by: Wang, Tevin, et al.
Published: (2025)
by: Wang, Tevin, et al.
Published: (2025)
Lugha-Llama: Adapting Large Language Models for African Languages
by: Buzaaba, Happy, et al.
Published: (2025)
by: Buzaaba, Happy, et al.
Published: (2025)
Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
by: Chen, Howard, et al.
Published: (2025)
by: Chen, Howard, et al.
Published: (2025)
Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training
by: Zhong, Zexuan, et al.
Published: (2024)
by: Zhong, Zexuan, et al.
Published: (2024)
SAP: Syntactic Attention Pruning for Transformer-based Language Models
by: Lee, Tzu-Yun, et al.
Published: (2025)
by: Lee, Tzu-Yun, et al.
Published: (2025)
OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention
by: Zhou, Chenyu, et al.
Published: (2026)
by: Zhou, Chenyu, et al.
Published: (2026)
AMPLIFY:Attention-based Mixup for Performance Improvement and Label Smoothing in Transformer
by: Yang, Leixin, et al.
Published: (2023)
by: Yang, Leixin, et al.
Published: (2023)
Short Data, Long Context: Distilling Positional Knowledge in Transformers
by: Huber, Patrick, et al.
Published: (2026)
by: Huber, Patrick, et al.
Published: (2026)
Parallelizing Linear Transformers with the Delta Rule over Sequence Length
by: Yang, Songlin, et al.
Published: (2024)
by: Yang, Songlin, et al.
Published: (2024)
LASER: Attention with Exponential Transformation
by: Duvvuri, Sai Surya, et al.
Published: (2024)
by: Duvvuri, Sai Surya, et al.
Published: (2024)
Weights to Code: Extracting Interpretable Algorithms from the Discrete Transformer
by: Zhang, Yifan, et al.
Published: (2026)
by: Zhang, Yifan, et al.
Published: (2026)
STENCIL: Submodular Mutual Information Based Weak Supervision for Cold-Start Active Learning
by: Beck, Nathan, et al.
Published: (2024)
by: Beck, Nathan, et al.
Published: (2024)
Position-Agnostic Pre-Projection for Transformer Attention: Nonlinear Feature Construction and Content Skip Before Q/K/V
by: Shinde, Chirag
Published: (2026)
by: Shinde, Chirag
Published: (2026)
The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
by: Zhu, Xinyu, et al.
Published: (2025)
by: Zhu, Xinyu, et al.
Published: (2025)
Faster Transformer Decoding: N-gram Masked Self-Attention
by: Chelba, Ciprian, et al.
Published: (2020)
by: Chelba, Ciprian, et al.
Published: (2020)
Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
by: Xia, Mengzhou, et al.
Published: (2023)
by: Xia, Mengzhou, et al.
Published: (2023)
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
by: Leemann, Tobias, et al.
Published: (2024)
by: Leemann, Tobias, et al.
Published: (2024)
Evaluating Large Language Models at Evaluating Instruction Following
by: Zeng, Zhiyuan, et al.
Published: (2023)
by: Zeng, Zhiyuan, et al.
Published: (2023)
Extracting Protein-Protein Interactions (PPIs) from Biomedical Literature using Attention-based Relational Context Information
by: Park, Gilchan, et al.
Published: (2024)
by: Park, Gilchan, et al.
Published: (2024)
RuleR: Improving LLM Controllability by Rule-based Data Recycling
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
Generalized Probabilistic Attention Mechanism in Transformers
by: Heo, DongNyeong, et al.
Published: (2024)
by: Heo, DongNyeong, et al.
Published: (2024)
Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
by: He, Zhengfu, et al.
Published: (2024)
by: He, Zhengfu, et al.
Published: (2024)
Selective Attention: Enhancing Transformer through Principled Context Control
by: Zhang, Xuechen, et al.
Published: (2024)
by: Zhang, Xuechen, et al.
Published: (2024)
Rethinking Attention Output Projection: Structured Hadamard Transforms for Efficient Transformers
by: Aggarwal, Shubham, et al.
Published: (2026)
by: Aggarwal, Shubham, et al.
Published: (2026)
Can LLMs Follow Simple Rules?
by: Mu, Norman, et al.
Published: (2023)
by: Mu, Norman, et al.
Published: (2023)
Organize the Web: Constructing Domains Enhances Pre-Training Data Curation
by: Wettig, Alexander, et al.
Published: (2025)
by: Wettig, Alexander, et al.
Published: (2025)
Metadata Conditioning Accelerates Language Model Pre-training
by: Gao, Tianyu, et al.
Published: (2025)
by: Gao, Tianyu, et al.
Published: (2025)
Similar Items
-
Finding Transformer Circuits with Edge Pruning
by: Bhaskar, Adithya, et al.
Published: (2024) -
The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models
by: Bhaskar, Adithya, et al.
Published: (2024) -
Representing Rule-based Chatbots with Transformers
by: Friedman, Dan, et al.
Published: (2024) -
QuRating: Selecting High-Quality Data for Training Language Models
by: Wettig, Alexander, et al.
Published: (2024) -
How to Train Long-Context Language Models (Effectively)
by: Gao, Tianyu, et al.
Published: (2024)