Pay Attention to Small Weights
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Chao, Jacobs, Tom, Gadhikar, Advait, Burkholz, Rebekka |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sign-In to the Lottery: Reparameterizing Sparse Training From Scratch
by: Gadhikar, Advait, et al.
Published: (2025)
by: Gadhikar, Advait, et al.
Published: (2025)
Masks, Signs, And Learning Rate Rewinding
by: Gadhikar, Advait, et al.
Published: (2024)
by: Gadhikar, Advait, et al.
Published: (2024)
Hyperbolic Aware Minimization: Implicit Bias for Sparsity
by: Jacobs, Tom, et al.
Published: (2025)
by: Jacobs, Tom, et al.
Published: (2025)
Cyclic Sparse Training: Is it Enough?
by: Gadhikar, Advait, et al.
Published: (2024)
by: Gadhikar, Advait, et al.
Published: (2024)
OptRot: Mitigating Weight Outliers via Data-Free Rotations for Post-Training Quantization
by: Gadhikar, Advait, et al.
Published: (2025)
by: Gadhikar, Advait, et al.
Published: (2025)
The Graphon Limit Hypothesis: Understanding Neural Network Pruning via Infinite Width Analysis
by: Pham, Hoang, et al.
Published: (2025)
by: Pham, Hoang, et al.
Published: (2025)
Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias?
by: Jacobs, Tom, et al.
Published: (2025)
by: Jacobs, Tom, et al.
Published: (2025)
Never Saddle for Reparameterized Steepest Descent as Mirror Flow
by: Jacobs, Tom, et al.
Published: (2026)
by: Jacobs, Tom, et al.
Published: (2026)
Mask in the Mirror: Implicit Sparsification
by: Jacobs, Tom, et al.
Published: (2024)
by: Jacobs, Tom, et al.
Published: (2024)
HORST: Composing Optimizer Geometries for Sparse Transformer Training
by: Jacobs, Tom, et al.
Published: (2026)
by: Jacobs, Tom, et al.
Published: (2026)
Bridging Domains through Subspace-Aware Model Merging
by: Chaves, Levy, et al.
Published: (2026)
by: Chaves, Levy, et al.
Published: (2026)
GATE: How to Keep Out Intrusive Neighbors
by: Mustafa, Nimrah, et al.
Published: (2024)
by: Mustafa, Nimrah, et al.
Published: (2024)
Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
by: De Muri, Giovanni, et al.
Published: (2025)
by: De Muri, Giovanni, et al.
Published: (2025)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
by: Collins, Liam, et al.
Published: (2024)
by: Collins, Liam, et al.
Published: (2024)
Fixed Aggregation Features Can Rival GNNs
by: Rubio-Madrigal, Celia, et al.
Published: (2026)
by: Rubio-Madrigal, Celia, et al.
Published: (2026)
Unmask It! AI-Generated Product Review Detection in Dravidian Languages
by: De, Somsubhra, et al.
Published: (2025)
by: De, Somsubhra, et al.
Published: (2025)
CAVACHON: a hierarchical variational autoencoder to integrate multi-modal single-cell data
by: Hsieh, Ping-Han, et al.
Published: (2024)
by: Hsieh, Ping-Han, et al.
Published: (2024)
Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation
by: Racioppo, Peter
Published: (2025)
by: Racioppo, Peter
Published: (2025)
Efficient Approximate Posterior Sampling with Annealed Langevin Monte Carlo
by: Parulekar, Advait, et al.
Published: (2025)
by: Parulekar, Advait, et al.
Published: (2025)
Weighted Graph Structure Learning with Attention Denoising for Node Classification
by: Wang, Tingting, et al.
Published: (2025)
by: Wang, Tingting, et al.
Published: (2025)
Robustness of Mixtures of Experts to Feature Noise
by: Sun, Dong, et al.
Published: (2026)
by: Sun, Dong, et al.
Published: (2026)
Learning to Pay Attention: Unsupervised Modeling of Attentive and Inattentive Respondents in Survey Data
by: Triantafyllopoulos, Ilias, et al.
Published: (2026)
by: Triantafyllopoulos, Ilias, et al.
Published: (2026)
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
by: Zheng, Haizhong, et al.
Published: (2025)
by: Zheng, Haizhong, et al.
Published: (2025)
More Expressive Attention with Negative Weights
by: Lv, Ang, et al.
Published: (2024)
by: Lv, Ang, et al.
Published: (2024)
SparseOpt: Addressing Normalization-induced Gradient Skew in Sparse Training
by: Adnan, Mohammed, et al.
Published: (2026)
by: Adnan, Mohammed, et al.
Published: (2026)
Nautile-370M: Spectral Memory Meets Attention in a Small Reasoning Model
by: Chenebaux, Maixent
Published: (2026)
by: Chenebaux, Maixent
Published: (2026)
Spend Search Where It Pays: Value-Guided Structured Sampling and Optimization for Generative Recommendation
by: Jiang, Jie, et al.
Published: (2026)
by: Jiang, Jie, et al.
Published: (2026)
AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions
by: Liu, Xianyang, et al.
Published: (2026)
by: Liu, Xianyang, et al.
Published: (2026)
LORA-CRAFT: Cross-layer Rank Adaptation via Frozen Tucker Decomposition of Pre-trained Attention Weights
by: Dewage, Kasun, et al.
Published: (2026)
by: Dewage, Kasun, et al.
Published: (2026)
Adaptive Dual-Weighting Framework for Federated Learning via Out-of-Distribution Detection
by: Ling, Zhiwei, et al.
Published: (2026)
by: Ling, Zhiwei, et al.
Published: (2026)
Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
by: Yin, Lu, et al.
Published: (2023)
by: Yin, Lu, et al.
Published: (2023)
Progressive Sparse Attention: Algorithm and System Co-design for Efficient Attention in LLM Serving
by: Zhou, Qihui, et al.
Published: (2025)
by: Zhou, Qihui, et al.
Published: (2025)
Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM Agents
by: Liu, Zhihan, et al.
Published: (2026)
by: Liu, Zhihan, et al.
Published: (2026)
Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers
by: Karbevski, Marko, et al.
Published: (2025)
by: Karbevski, Marko, et al.
Published: (2025)
Multi-Granular Attention based Heterogeneous Hypergraph Neural Network
by: Jin, Hong, et al.
Published: (2025)
by: Jin, Hong, et al.
Published: (2025)
Spectral Graph Pruning Against Over-Squashing and Over-Smoothing
by: Jamadandi, Adarsh, et al.
Published: (2024)
by: Jamadandi, Adarsh, et al.
Published: (2024)
Pay Less Attention to Deceptive Artifacts: Robust Detection of Compressed Deepfakes on Online Social Networks
by: Li, Manyi, et al.
Published: (2025)
by: Li, Manyi, et al.
Published: (2025)
MultiMax: Sparse and Multi-Modal Attention Learning
by: Zhou, Yuxuan, et al.
Published: (2024)
by: Zhou, Yuxuan, et al.
Published: (2024)
Accelerating PayPal's Commerce Agent with Speculative Decoding: An Empirical Study on EAGLE3 with Fine-Tuned Nemotron Models
by: Qin, Ally, et al.
Published: (2026)
by: Qin, Ally, et al.
Published: (2026)
Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformers
by: Zhang, Yukun, et al.
Published: (2025)
by: Zhang, Yukun, et al.
Published: (2025)
Similar Items
-
Sign-In to the Lottery: Reparameterizing Sparse Training From Scratch
by: Gadhikar, Advait, et al.
Published: (2025) -
Masks, Signs, And Learning Rate Rewinding
by: Gadhikar, Advait, et al.
Published: (2024) -
Hyperbolic Aware Minimization: Implicit Bias for Sparsity
by: Jacobs, Tom, et al.
Published: (2025) -
Cyclic Sparse Training: Is it Enough?
by: Gadhikar, Advait, et al.
Published: (2024) -
OptRot: Mitigating Weight Outliers via Data-Free Rotations for Post-Training Quantization
by: Gadhikar, Advait, et al.
Published: (2025)