The Quantization Benefits of Residual-Free Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ji, Yiping, Sabanayagam, Mahalakshmi, Moghadam, Peyman, Saratchandran, Hemanth, Lucey, Simon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Always Skip Attention
von: Ji, Yiping, et al.
Veröffentlicht: (2025)
von: Ji, Yiping, et al.
Veröffentlicht: (2025)
Cutting the Skip: Training Residual-Free Transformers
von: Ji, Yiping, et al.
Veröffentlicht: (2025)
von: Ji, Yiping, et al.
Veröffentlicht: (2025)
Spectral Conditioning of Attention Improves Transformer Performance
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2026)
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2026)
Rethinking Attention: Polynomial Alternatives to Softmax in Transformers
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2024)
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2024)
Leaner Transformers: More Heads, Less Depth
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2025)
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2025)
SineLoRA$Δ$: Sine-Activated Delta Compression
von: Gordon, Cameron, et al.
Veröffentlicht: (2025)
von: Gordon, Cameron, et al.
Veröffentlicht: (2025)
The Inlet Rank Collapse in Implicit Neural Representations: Diagnosis and Unified Remedy
von: Zheng, Jianqiao, et al.
Veröffentlicht: (2026)
von: Zheng, Jianqiao, et al.
Veröffentlicht: (2026)
From Tables to Signals: Revealing Spectral Adaptivity in TabPFN
von: Zheng, Jianqiao, et al.
Veröffentlicht: (2025)
von: Zheng, Jianqiao, et al.
Veröffentlicht: (2025)
Efficient Learning With Sine-Activated Low-rank Matrices
von: Ji, Yiping, et al.
Veröffentlicht: (2024)
von: Ji, Yiping, et al.
Veröffentlicht: (2024)
From Activation to Initialization: Scaling Insights for Optimizing Neural Fields
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2024)
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2024)
Analyzing the Neural Tangent Kernel of Periodically Activated Coordinate Networks
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2024)
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2024)
Architectural Strategies for the optimization of Physics-Informed Neural Networks
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2024)
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2024)
Preconditioned Attention: Enhancing Efficiency in Transformers
von: Saratchandran, Hemanth
Veröffentlicht: (2026)
von: Saratchandran, Hemanth
Veröffentlicht: (2026)
Preconditioners for the Stochastic Training of Neural Fields
von: Chng, Shin-Fang, et al.
Veröffentlicht: (2024)
von: Chng, Shin-Fang, et al.
Veröffentlicht: (2024)
Stable Forgetting: Bounded Parameter-Efficient Unlearning in Foundation Models
von: Garg, Arpit, et al.
Veröffentlicht: (2025)
von: Garg, Arpit, et al.
Veröffentlicht: (2025)
Enhancing Transformers Through Conditioned Embedded Tokens
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2025)
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2025)
A Sampling Theory Perspective on Activations for Implicit Neural Representations
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2024)
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2024)
D'OH: Decoder-Only Random Hypernetworks for Implicit Neural Representations
von: Gordon, Cameron, et al.
Veröffentlicht: (2024)
von: Gordon, Cameron, et al.
Veröffentlicht: (2024)
Exact Generalisation Error Exposes Benchmarks Skew Graph Neural Networks Success (or Failure)
von: Ayday, Nil, et al.
Veröffentlicht: (2025)
von: Ayday, Nil, et al.
Veröffentlicht: (2025)
Cluster Specific Representation Learning
von: Sabanayagam, Mahalakshmi, et al.
Veröffentlicht: (2024)
von: Sabanayagam, Mahalakshmi, et al.
Veröffentlicht: (2024)
Exact Certification of (Graph) Neural Networks Against Label Poisoning
von: Sabanayagam, Mahalakshmi, et al.
Veröffentlicht: (2024)
von: Sabanayagam, Mahalakshmi, et al.
Veröffentlicht: (2024)
Robustness Certificates for Neural Networks against Adversarial Attacks
von: Taheri, Sara, et al.
Veröffentlicht: (2025)
von: Taheri, Sara, et al.
Veröffentlicht: (2025)
Provable Robustness of (Graph) Neural Networks Against Data Poisoning and Backdoor Attacks
von: Gosch, Lukas, et al.
Veröffentlicht: (2024)
von: Gosch, Lukas, et al.
Veröffentlicht: (2024)
Robust Feature Inference: A Test-time Defense Strategy using Spectral Projections
von: Singh, Anurag, et al.
Veröffentlicht: (2023)
von: Singh, Anurag, et al.
Veröffentlicht: (2023)
Generalization Certificates for Adversarially Robust Bayesian Linear Regression
von: Sabanayagam, Mahalakshmi, et al.
Veröffentlicht: (2025)
von: Sabanayagam, Mahalakshmi, et al.
Veröffentlicht: (2025)
Different Statistical Perspectives for Understanding Generalisation in Graph Neural Networks
von: Ayday, Nil, et al.
Veröffentlicht: (2026)
von: Ayday, Nil, et al.
Veröffentlicht: (2026)
Structured Initialization for Vision Transformers
von: Zheng, Jianqiao, et al.
Veröffentlicht: (2025)
von: Zheng, Jianqiao, et al.
Veröffentlicht: (2025)
SineProject: Machine Unlearning for Stable Vision Language Alignment
von: Garg, Arpit, et al.
Veröffentlicht: (2025)
von: Garg, Arpit, et al.
Veröffentlicht: (2025)
Exact Certification of Neural Networks and Partition Aggregation Ensembles against Label Poisoning
von: Mohgaonkar, Ajinkya, et al.
Veröffentlicht: (2026)
von: Mohgaonkar, Ajinkya, et al.
Veröffentlicht: (2026)
Weight Conditioning for Smooth Optimization of Neural Networks
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2024)
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2024)
Transformers Pretrained on Procedural Data Contain Modular Structures for Algorithmic Reasoning
von: Shinnick, Zachary, et al.
Veröffentlicht: (2025)
von: Shinnick, Zachary, et al.
Veröffentlicht: (2025)
Data Denoising and Derivative Estimation for Data-Driven Modeling of Nonlinear Dynamical Systems
von: Yao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Yao, Jiaqi, et al.
Veröffentlicht: (2025)
Mask the Target: A Plug-and-Play Regularizer Against LoRA Forgetting
von: Xu, Runze, et al.
Veröffentlicht: (2026)
von: Xu, Runze, et al.
Veröffentlicht: (2026)
Invertible Neural Warp for NeRF
von: Chng, Shin-Fang, et al.
Veröffentlicht: (2024)
von: Chng, Shin-Fang, et al.
Veröffentlicht: (2024)
Gradient Descent as a Shrinkage Operator for Spectral Bias
von: Lucey, Simon
Veröffentlicht: (2025)
von: Lucey, Simon
Veröffentlicht: (2025)
Spatioformer: A Geo-encoded Transformer for Large-Scale Plant Species Richness Prediction
von: Guo, Yiqing, et al.
Veröffentlicht: (2024)
von: Guo, Yiqing, et al.
Veröffentlicht: (2024)
Procedural Pretraining: Warming Up Language Models with Abstract Data
von: Jiang, Liangze, et al.
Veröffentlicht: (2026)
von: Jiang, Liangze, et al.
Veröffentlicht: (2026)
GoQuant: Geometric Orthogonal Residual Projection for Multiplier-Free Power-of-Two Transformer Quantization
von: Xiang, Maoyang, et al.
Veröffentlicht: (2026)
von: Xiang, Maoyang, et al.
Veröffentlicht: (2026)
Flashbacks to Harmonize Stability and Plasticity in Continual Learning
von: Mahmoodi, Leila, et al.
Veröffentlicht: (2025)
von: Mahmoodi, Leila, et al.
Veröffentlicht: (2025)
Quantization-Free Autoregressive Action Transformer
von: Sheebaelhamd, Ziyad, et al.
Veröffentlicht: (2025)
von: Sheebaelhamd, Ziyad, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Always Skip Attention
von: Ji, Yiping, et al.
Veröffentlicht: (2025) -
Cutting the Skip: Training Residual-Free Transformers
von: Ji, Yiping, et al.
Veröffentlicht: (2025) -
Spectral Conditioning of Attention Improves Transformer Performance
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2026) -
Rethinking Attention: Polynomial Alternatives to Softmax in Transformers
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2024) -
Leaner Transformers: More Heads, Less Depth
von: Saratchandran, Hemanth, et al.
Veröffentlicht: (2025)