Post-Norm can Resharpen Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zsámboki, Pál, Levi, Benjamin, Smith, David Ansel Josef, Kagalwala, Mitansh, Kell, Arlington, Liechty, Samuel, Wang, Cong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Do Attention Heads Compete or Cooperate during Counting?
von: Zsámboki, Pál, et al.
Veröffentlicht: (2025)
von: Zsámboki, Pál, et al.
Veröffentlicht: (2025)
Decoupled Weight Decay for Any $p$ Norm
von: Outmezguine, Nadav Joseph, et al.
Veröffentlicht: (2024)
von: Outmezguine, Nadav Joseph, et al.
Veröffentlicht: (2024)
Trust by Design: Skill Profiles for Transparent, Cost-Aware LLM Routing
von: Okamoto, Mika, et al.
Veröffentlicht: (2026)
von: Okamoto, Mika, et al.
Veröffentlicht: (2026)
Attend or Perish: Benchmarking Attention in Algorithmic Reasoning
von: Spiegel, Michal, et al.
Veröffentlicht: (2025)
von: Spiegel, Michal, et al.
Veröffentlicht: (2025)
Are Graph Attention Networks Able to Model Structural Information?
von: Noravesh, Farshad, et al.
Veröffentlicht: (2025)
von: Noravesh, Farshad, et al.
Veröffentlicht: (2025)
Intrinsically Interpretable Attention via Sparse Post-Training
von: Draye, Florent, et al.
Veröffentlicht: (2025)
von: Draye, Florent, et al.
Veröffentlicht: (2025)
Optimal Scaling Needs Optimal Norm
von: Filatov, Oleg, et al.
Veröffentlicht: (2025)
von: Filatov, Oleg, et al.
Veröffentlicht: (2025)
MuaLLM: A Multimodal Large Language Model Agent for Circuit Design Assistance with Hybrid Contextual Retrieval-Augmented Generation
von: Abbineni, Pravallika, et al.
Veröffentlicht: (2025)
von: Abbineni, Pravallika, et al.
Veröffentlicht: (2025)
GradientStabilizer:Fix the Norm, Not the Gradient
von: Huang, Tianjin, et al.
Veröffentlicht: (2025)
von: Huang, Tianjin, et al.
Veröffentlicht: (2025)
BoA: Attention-aware Post-training Quantization without Backpropagation
von: Kim, Junhan, et al.
Veröffentlicht: (2024)
von: Kim, Junhan, et al.
Veröffentlicht: (2024)
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
von: Chen, Shimao, et al.
Veröffentlicht: (2024)
von: Chen, Shimao, et al.
Veröffentlicht: (2024)
Verifier-Free RL for LLMs via Intrinsic Gradient-Norm Reward
von: Wen, Xuexiang, et al.
Veröffentlicht: (2026)
von: Wen, Xuexiang, et al.
Veröffentlicht: (2026)
VL Norm: Rethink Loss Aggregation in RLVR
von: He, Zhiyuan, et al.
Veröffentlicht: (2025)
von: He, Zhiyuan, et al.
Veröffentlicht: (2025)
Norm Anchors Make Model Edits Last
von: Liu, Mingda, et al.
Veröffentlicht: (2026)
von: Liu, Mingda, et al.
Veröffentlicht: (2026)
Post-Training Sparse Attention with Double Sparsity
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
Beyond the Norms: Detecting Prediction Errors in Regression Models
von: Altieri, Andres, et al.
Veröffentlicht: (2024)
von: Altieri, Andres, et al.
Veröffentlicht: (2024)
The Geometry of Grokking: Norm Minimization on the Zero-Loss Manifold
von: Musat, Tiberiu
Veröffentlicht: (2025)
von: Musat, Tiberiu
Veröffentlicht: (2025)
Learning Topological Representations with Bidirectional Graph Attention Network for Solving Job Shop Scheduling Problem
von: Zhang, Cong, et al.
Veröffentlicht: (2024)
von: Zhang, Cong, et al.
Veröffentlicht: (2024)
The Ky Fan Norms and Beyond: Dual Norms and Combinations for Matrix Optimization
von: Kravatskiy, Alexey, et al.
Veröffentlicht: (2025)
von: Kravatskiy, Alexey, et al.
Veröffentlicht: (2025)
Recency Biased Causal Attention for Time-series Forecasting
von: Hegazy, Kareem, et al.
Veröffentlicht: (2025)
von: Hegazy, Kareem, et al.
Veröffentlicht: (2025)
ORLoopBench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality in Operations Research
von: Ao, Ruicheng, et al.
Veröffentlicht: (2026)
von: Ao, Ruicheng, et al.
Veröffentlicht: (2026)
OptiRepair: Closed-Loop Diagnosis and Repair of Supply Chain Optimization Models with LLM Agents
von: Ao, Ruicheng, et al.
Veröffentlicht: (2026)
von: Ao, Ruicheng, et al.
Veröffentlicht: (2026)
ConjNorm: Tractable Density Estimation for Out-of-Distribution Detection
von: Peng, Bo, et al.
Veröffentlicht: (2024)
von: Peng, Bo, et al.
Veröffentlicht: (2024)
Unveiling Options with Neural Decomposition
von: Alikhasi, Mahdi, et al.
Veröffentlicht: (2024)
von: Alikhasi, Mahdi, et al.
Veröffentlicht: (2024)
Aligned, Orthogonal or In-conflict: When can we safely optimize Chain-of-Thought?
von: Kaufmann, Max, et al.
Veröffentlicht: (2026)
von: Kaufmann, Max, et al.
Veröffentlicht: (2026)
Exploring Sparsity and Smoothness of Arbitrary $\ell_p$ Norms in Adversarial Attacks
von: Duhme, Christof, et al.
Veröffentlicht: (2026)
von: Duhme, Christof, et al.
Veröffentlicht: (2026)
Unpacking the Implicit Norm Dynamics of Sharpness-Aware Minimization in Tensorized Models
von: Cao, Tianxiao, et al.
Veröffentlicht: (2025)
von: Cao, Tianxiao, et al.
Veröffentlicht: (2025)
Selective Steering: Norm-Preserving Control Through Discriminative Layer Selection
von: Dang, Quy-Anh, et al.
Veröffentlicht: (2026)
von: Dang, Quy-Anh, et al.
Veröffentlicht: (2026)
A Simple Model of Inference Scaling Laws
von: Levi, Noam
Veröffentlicht: (2024)
von: Levi, Noam
Veröffentlicht: (2024)
Hard Examples Are All You Need: Maximizing GRPO Post-Training Under Annotation Budgets
von: Pikus, Benjamin, et al.
Veröffentlicht: (2025)
von: Pikus, Benjamin, et al.
Veröffentlicht: (2025)
Measuring Social Norms of Large Language Models
von: Yuan, Ye, et al.
Veröffentlicht: (2024)
von: Yuan, Ye, et al.
Veröffentlicht: (2024)
Local vs Global continual learning
von: Lanzillotta, Giulia, et al.
Veröffentlicht: (2024)
von: Lanzillotta, Giulia, et al.
Veröffentlicht: (2024)
LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
von: Helff, Lukas, et al.
Veröffentlicht: (2026)
von: Helff, Lukas, et al.
Veröffentlicht: (2026)
Classifying Overlapping Gaussian Mixtures in High Dimensions: From Optimal Classifiers to Neural Nets
von: Cohen, Khen, et al.
Veröffentlicht: (2024)
von: Cohen, Khen, et al.
Veröffentlicht: (2024)
Attention Sinks and Outliers in Attention Residuals
von: Luo, Haozheng, et al.
Veröffentlicht: (2026)
von: Luo, Haozheng, et al.
Veröffentlicht: (2026)
Sufficient Conditions for Stability of Minimum-Norm Interpolating Deep ReLU Networks
von: Harzli, Ouns El, et al.
Veröffentlicht: (2026)
von: Harzli, Ouns El, et al.
Veröffentlicht: (2026)
AtMan: Understanding Transformer Predictions Through Memory Efficient Attention Manipulation
von: Deiseroth, Björn, et al.
Veröffentlicht: (2023)
von: Deiseroth, Björn, et al.
Veröffentlicht: (2023)
In-Context Learning can Perform Continual Learning Like Humans
von: Kang, Liuwang, et al.
Veröffentlicht: (2025)
von: Kang, Liuwang, et al.
Veröffentlicht: (2025)
Transformers can do Bayesian Clustering
von: Bhaskaran, Prajit, et al.
Veröffentlicht: (2025)
von: Bhaskaran, Prajit, et al.
Veröffentlicht: (2025)
Graph Diffusion that can Insert and Delete
von: Ninniri, Matteo, et al.
Veröffentlicht: (2025)
von: Ninniri, Matteo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Do Attention Heads Compete or Cooperate during Counting?
von: Zsámboki, Pál, et al.
Veröffentlicht: (2025) -
Decoupled Weight Decay for Any $p$ Norm
von: Outmezguine, Nadav Joseph, et al.
Veröffentlicht: (2024) -
Trust by Design: Skill Profiles for Transparent, Cost-Aware LLM Routing
von: Okamoto, Mika, et al.
Veröffentlicht: (2026) -
Attend or Perish: Benchmarking Attention in Algorithmic Reasoning
von: Spiegel, Michal, et al.
Veröffentlicht: (2025) -
Are Graph Attention Networks Able to Model Structural Information?
von: Noravesh, Farshad, et al.
Veröffentlicht: (2025)