Reversed Attention: On The Gradient Descent Of Attention Layers In GPT
Fuente:
arXiv
Saved in:
| Main Authors: | Katz, Shahar, Wolf, Lior |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Segment-Based Attention Masking for GPTs
by: Katz, Shahar, et al.
Published: (2024)
by: Katz, Shahar, et al.
Published: (2024)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
by: Katz, Shahar, et al.
Published: (2024)
by: Katz, Shahar, et al.
Published: (2024)
Detecting and Pruning Prominent but Detrimental Neurons in Large Language Models
by: Ali, Ameen, et al.
Published: (2025)
by: Ali, Ameen, et al.
Published: (2025)
TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors
by: Atad, Ido Andrew, et al.
Published: (2026)
by: Atad, Ido Andrew, et al.
Published: (2026)
Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers
by: Wong, Liang Ze
Published: (2025)
by: Wong, Liang Ze
Published: (2025)
PRILoRA: Pruned and Rank-Increasing Low-Rank Adaptation
by: Benedek, Nadav, et al.
Published: (2024)
by: Benedek, Nadav, et al.
Published: (2024)
AlignTree: Efficient Defense Against LLM Jailbreak Attacks
by: Goren, Gil, et al.
Published: (2025)
by: Goren, Gil, et al.
Published: (2025)
Execution Guided Line-by-Line Code Generation
by: Lavon, Boaz, et al.
Published: (2025)
by: Lavon, Boaz, et al.
Published: (2025)
ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
by: Lior, Gili, et al.
Published: (2025)
by: Lior, Gili, et al.
Published: (2025)
High-Layer Attention Pruning with Rescaling
by: Liu, Songtao, et al.
Published: (2025)
by: Liu, Songtao, et al.
Published: (2025)
LayerBoost: Layer-Aware Attention Reduction for Efficient LLMs
by: Souibgui, Mohamed Ali, et al.
Published: (2026)
by: Souibgui, Mohamed Ali, et al.
Published: (2026)
Diffusion-Based Attention Warping for Consistent 3D Scene Editing
by: Gomel, Eyal, et al.
Published: (2024)
by: Gomel, Eyal, et al.
Published: (2024)
Multi-Layer Attention is the Amplifier of Demonstration Effectiveness
by: Wang, Dingzirui, et al.
Published: (2025)
by: Wang, Dingzirui, et al.
Published: (2025)
Skip-Layer Attention: Bridging Abstract and Detailed Dependencies in Transformers
by: Chen, Qian, et al.
Published: (2024)
by: Chen, Qian, et al.
Published: (2024)
Optimizing Attention with Mirror Descent: Generalized Max-Margin Token Selection
by: Julistiono, Addison Kristanto, et al.
Published: (2024)
by: Julistiono, Addison Kristanto, et al.
Published: (2024)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
by: Bozic, Vukasin, et al.
Published: (2023)
by: Bozic, Vukasin, et al.
Published: (2023)
SEAM: A Stochastic Benchmark for Multi-Document Tasks
by: Lior, Gili, et al.
Published: (2024)
by: Lior, Gili, et al.
Published: (2024)
Attend First, Consolidate Later: On the Importance of Attention in Different LLM Layers
by: Ben-Artzy, Amit, et al.
Published: (2024)
by: Ben-Artzy, Amit, et al.
Published: (2024)
Learning to Focus: Causal Attention Distillation via Gradient-Guided Token Pruning
by: Guo, Yiju, et al.
Published: (2025)
by: Guo, Yiju, et al.
Published: (2025)
Mitigating Copy Bias in In-Context Learning through Neuron Pruning
by: Ali, Ameen, et al.
Published: (2024)
by: Ali, Ameen, et al.
Published: (2024)
Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance
by: Alon, Bar, et al.
Published: (2026)
by: Alon, Bar, et al.
Published: (2026)
Can Gradient Descent Simulate Prompting?
by: Zhang, Eric, et al.
Published: (2025)
by: Zhang, Eric, et al.
Published: (2025)
In-context Learning and Gradient Descent Revisited
by: Deutch, Gilad, et al.
Published: (2023)
by: Deutch, Gilad, et al.
Published: (2023)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
by: Fu, Zizhuo, et al.
Published: (2026)
by: Fu, Zizhuo, et al.
Published: (2026)
Steered Generation via Gradient Descent on Sparse Features
by: Bhattacharyya, Sumanta, et al.
Published: (2025)
by: Bhattacharyya, Sumanta, et al.
Published: (2025)
HyLRA: Hybrid Layer Reuse Attention for Efficient Long-Context Inference
by: Ai, Xuan, et al.
Published: (2026)
by: Ai, Xuan, et al.
Published: (2026)
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
by: Musat, Tiberiu
Published: (2024)
by: Musat, Tiberiu
Published: (2024)
CLAA: Cross-Layer Attention Aggregation for Accelerating LLM Prefill
by: McDanel, Bradley, et al.
Published: (2026)
by: McDanel, Bradley, et al.
Published: (2026)
Cross-Layer Attention Probing for Fine-Grained Hallucination Detection
by: Suresh, Malavika, et al.
Published: (2025)
by: Suresh, Malavika, et al.
Published: (2025)
Introducing MAPO: Momentum-Aided Gradient Descent Prompt Optimization
by: Cui, Anthony, et al.
Published: (2024)
by: Cui, Anthony, et al.
Published: (2024)
Attention Instruction: Amplifying Attention in the Middle via Prompting
by: Zhang, Meiru, et al.
Published: (2024)
by: Zhang, Meiru, et al.
Published: (2024)
Adaptive Integrated Layered Attention (AILA)
by: Claster, William, et al.
Published: (2025)
by: Claster, William, et al.
Published: (2025)
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity
by: Ma, Da, et al.
Published: (2024)
by: Ma, Da, et al.
Published: (2024)
IlluSign: Illustrating Sign Language Videos by Leveraging the Attention Mechanism
by: Bruner, Janna, et al.
Published: (2025)
by: Bruner, Janna, et al.
Published: (2025)
Attention Residuals
by: Kimi Team, et al.
Published: (2026)
by: Kimi Team, et al.
Published: (2026)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
by: Brandon, William, et al.
Published: (2024)
by: Brandon, William, et al.
Published: (2024)
Distilling to Hybrid Attention Models via KL-Guided Layer Selection
by: Li, Yanhong, et al.
Published: (2025)
by: Li, Yanhong, et al.
Published: (2025)
Wave-PDE Nets: Trainable Wave-Equation Layers as an Alternative to Attention
by: Vejendla, Harshil
Published: (2025)
by: Vejendla, Harshil
Published: (2025)
SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs
by: Gao, Yizhao, et al.
Published: (2024)
by: Gao, Yizhao, et al.
Published: (2024)
Simulating Hard Attention Using Soft Attention
by: Yang, Andy, et al.
Published: (2024)
by: Yang, Andy, et al.
Published: (2024)
Similar Items
-
Segment-Based Attention Masking for GPTs
by: Katz, Shahar, et al.
Published: (2024) -
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
by: Katz, Shahar, et al.
Published: (2024) -
Detecting and Pruning Prominent but Detrimental Neurons in Large Language Models
by: Ali, Ameen, et al.
Published: (2025) -
TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors
by: Atad, Ido Andrew, et al.
Published: (2026) -
Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers
by: Wong, Liang Ze
Published: (2025)