Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
Fuente:
arXiv
Saved in:
| Main Authors: | Zuhri, Zayd M. K., Fuadi, Erland Hilman, Aji, Alham Fikri |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Predicting the Order of Upcoming Tokens Improves Language Modeling
by: Zuhri, Zayd M. K., et al.
Published: (2025)
by: Zuhri, Zayd M. K., et al.
Published: (2025)
MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding
by: Zuhri, Zayd Muhammad Kawakibi, et al.
Published: (2024)
by: Zuhri, Zayd Muhammad Kawakibi, et al.
Published: (2024)
QLESS: A Quantized Approach for Data Valuation and Selection in Large Language Model Fine-Tuning
by: Ananta, Moses, et al.
Published: (2025)
by: Ananta, Moses, et al.
Published: (2025)
LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation
by: Irawan, Patrick Amadeus, et al.
Published: (2026)
by: Irawan, Patrick Amadeus, et al.
Published: (2026)
COPAL-ID: Indonesian Language Reasoning with Local Culture and Nuances
by: Wibowo, Haryo Akbarianto, et al.
Published: (2023)
by: Wibowo, Haryo Akbarianto, et al.
Published: (2023)
COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training
by: Sakip, Akhmed, et al.
Published: (2026)
by: Sakip, Akhmed, et al.
Published: (2026)
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
by: Chen, Yihong, et al.
Published: (2026)
by: Chen, Yihong, et al.
Published: (2026)
Beyond Transfer Accuracy: Faithful Circuits for Controlled Low-Resource Adaptation
by: Nur'aini, Khumaisa, et al.
Published: (2026)
by: Nur'aini, Khumaisa, et al.
Published: (2026)
Attention Sinks Are Provably Necessary in Softmax Transformers: Evidence from Trigger-Conditional Tasks
by: Ran-Milo, Yuval
Published: (2026)
by: Ran-Milo, Yuval
Published: (2026)
LoraxBench: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages
by: Aji, Alham Fikri, et al.
Published: (2025)
by: Aji, Alham Fikri, et al.
Published: (2025)
Improving Low-Resource Machine Translation via Round-Trip Reinforcement Learning
by: Attia, Ahmed, et al.
Published: (2026)
by: Attia, Ahmed, et al.
Published: (2026)
Daisy-TTS: Simulating Wider Spectrum of Emotions via Prosody Embedding Decomposition
by: Chevi, Rendi, et al.
Published: (2024)
by: Chevi, Rendi, et al.
Published: (2024)
Unveiling the Influence of Amplifying Language-Specific Neurons
by: Rahmanisa, Inaya, et al.
Published: (2025)
by: Rahmanisa, Inaya, et al.
Published: (2025)
NusaAksara: A Multimodal and Multilingual Benchmark for Preserving Indonesian Indigenous Scripts
by: Adilazuarda, Muhammad Farid, et al.
Published: (2025)
by: Adilazuarda, Muhammad Farid, et al.
Published: (2025)
Does Visual Rendering Bypass Tokenization? Investigating Script-Tokenizer Misalignment in Pixel-Based Language Models
by: Susanto, Lucky, et al.
Published: (2026)
by: Susanto, Lucky, et al.
Published: (2026)
CLIP meets DINO for Tuning Zero-Shot Classifier using Unlabeled Image Collections
by: Imam, Mohamed Fazli, et al.
Published: (2024)
by: Imam, Mohamed Fazli, et al.
Published: (2024)
Attention Sinks in Massively Multilingual Neural Machine Translation:Discovery, Analysis, and Mitigation
by: Mutisya, Hillary, et al.
Published: (2026)
by: Mutisya, Hillary, et al.
Published: (2026)
Scalable-Softmax Is Superior for Attention
by: Nakanishi, Ken M.
Published: (2025)
by: Nakanishi, Ken M.
Published: (2025)
Universal Approximation with Softmax Attention
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Provable Robustness Against a Union of $\ell_0$ Adversarial Attacks
by: Hammoudeh, Zayd, et al.
Published: (2023)
by: Hammoudeh, Zayd, et al.
Published: (2023)
Training Data Influence Analysis and Estimation: A Survey
by: Hammoudeh, Zayd, et al.
Published: (2022)
by: Hammoudeh, Zayd, et al.
Published: (2022)
Why Softmax Attention Outperforms Linear Attention
by: Deng, Yichuan, et al.
Published: (2023)
by: Deng, Yichuan, et al.
Published: (2023)
Attention Sinks and Outliers in Attention Residuals
by: Luo, Haozheng, et al.
Published: (2026)
by: Luo, Haozheng, et al.
Published: (2026)
ASAP: Attention Sink Anchored Pruning
by: Lee, Jaehyuk, et al.
Published: (2026)
by: Lee, Jaehyuk, et al.
Published: (2026)
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
by: Son, Seungwoo, et al.
Published: (2024)
by: Son, Seungwoo, et al.
Published: (2024)
On the Invariants of Softmax Attention
by: Lee, Wonsuk
Published: (2026)
by: Lee, Wonsuk
Published: (2026)
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
by: Nishikawa, Naoki, et al.
Published: (2025)
by: Nishikawa, Naoki, et al.
Published: (2025)
Softmax Attention with Constant Cost per Token
by: Heinsen, Franz A.
Published: (2024)
by: Heinsen, Franz A.
Published: (2024)
Balanced Multi-Factor In-Context Learning for Multilingual Large Language Models
by: Kaneko, Masahiro, et al.
Published: (2025)
by: Kaneko, Masahiro, et al.
Published: (2025)
Extracting General-use Transformers for Low-resource Languages via Knowledge Distillation
by: Cruz, Jan Christian Blaise, et al.
Published: (2025)
by: Cruz, Jan Christian Blaise, et al.
Published: (2025)
Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation
by: Mansurov, Jonibek, et al.
Published: (2024)
by: Mansurov, Jonibek, et al.
Published: (2024)
Sense Representations Are Inducible Interfaces
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)
LLM Olympiad: Why Model Evaluation Needs a Sealed Exam
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)
by: Cruz, Jan Christian Blaise, et al.
Published: (2026)
Beyond Probabilities: Unveiling the Misalignment in Evaluating Large Language Models
by: Lyu, Chenyang, et al.
Published: (2024)
by: Lyu, Chenyang, et al.
Published: (2024)
On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective
by: Mongaras, Gabriel, et al.
Published: (2025)
by: Mongaras, Gabriel, et al.
Published: (2025)
Customizing the Inductive Biases of Softmax Attention using Structured Matrices
by: Kuang, Yilun, et al.
Published: (2025)
by: Kuang, Yilun, et al.
Published: (2025)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
by: Fu, Zizhuo, et al.
Published: (2026)
by: Fu, Zizhuo, et al.
Published: (2026)
Stochastic Parroting in Temporal Attention -- Regulating the Diagonal Sink
by: Hankemeier, Victoria, et al.
Published: (2026)
by: Hankemeier, Victoria, et al.
Published: (2026)
The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry
by: Zhang, Michael, et al.
Published: (2024)
by: Zhang, Michael, et al.
Published: (2024)
Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning
by: Xie, Zixuan, et al.
Published: (2026)
by: Xie, Zixuan, et al.
Published: (2026)
Similar Items
-
Predicting the Order of Upcoming Tokens Improves Language Modeling
by: Zuhri, Zayd M. K., et al.
Published: (2025) -
MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding
by: Zuhri, Zayd Muhammad Kawakibi, et al.
Published: (2024) -
QLESS: A Quantized Approach for Data Valuation and Selection in Large Language Model Fine-Tuning
by: Ananta, Moses, et al.
Published: (2025) -
LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation
by: Irawan, Patrick Amadeus, et al.
Published: (2026) -
COPAL-ID: Indonesian Language Reasoning with Local Culture and Nuances
by: Wibowo, Haryo Akbarianto, et al.
Published: (2023)