Saved in:
| Main Authors: | Ji, Yiping, Martens, James, Zheng, Jianqiao, Zhou, Ziqin, Moghadam, Peyman, Zhang, Xinyu, Saratchandran, Hemanth, Lucey, Simon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2510.00345 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Always Skip Attention
by: Ji, Yiping, et al.
Published: (2025)
by: Ji, Yiping, et al.
Published: (2025)
The Quantization Benefits of Residual-Free Transformers
by: Ji, Yiping, et al.
Published: (2026)
by: Ji, Yiping, et al.
Published: (2026)
Rethinking Attention: Polynomial Alternatives to Softmax in Transformers
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
The Inlet Rank Collapse in Implicit Neural Representations: Diagnosis and Unified Remedy
by: Zheng, Jianqiao, et al.
Published: (2026)
by: Zheng, Jianqiao, et al.
Published: (2026)
From Tables to Signals: Revealing Spectral Adaptivity in TabPFN
by: Zheng, Jianqiao, et al.
Published: (2025)
by: Zheng, Jianqiao, et al.
Published: (2025)
Spectral Conditioning of Attention Improves Transformer Performance
by: Saratchandran, Hemanth, et al.
Published: (2026)
by: Saratchandran, Hemanth, et al.
Published: (2026)
Leaner Transformers: More Heads, Less Depth
by: Saratchandran, Hemanth, et al.
Published: (2025)
by: Saratchandran, Hemanth, et al.
Published: (2025)
SineLoRA$Δ$: Sine-Activated Delta Compression
by: Gordon, Cameron, et al.
Published: (2025)
by: Gordon, Cameron, et al.
Published: (2025)
Efficient Learning With Sine-Activated Low-rank Matrices
by: Ji, Yiping, et al.
Published: (2024)
by: Ji, Yiping, et al.
Published: (2024)
Preconditioners for the Stochastic Training of Neural Fields
by: Chng, Shin-Fang, et al.
Published: (2024)
by: Chng, Shin-Fang, et al.
Published: (2024)
From Activation to Initialization: Scaling Insights for Optimizing Neural Fields
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
Structured Initialization for Vision Transformers
by: Zheng, Jianqiao, et al.
Published: (2025)
by: Zheng, Jianqiao, et al.
Published: (2025)
Analyzing the Neural Tangent Kernel of Periodically Activated Coordinate Networks
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
Architectural Strategies for the optimization of Physics-Informed Neural Networks
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
Preconditioned Attention: Enhancing Efficiency in Transformers
by: Saratchandran, Hemanth
Published: (2026)
by: Saratchandran, Hemanth
Published: (2026)
Stable Forgetting: Bounded Parameter-Efficient Unlearning in Foundation Models
by: Garg, Arpit, et al.
Published: (2025)
by: Garg, Arpit, et al.
Published: (2025)
Enhancing Transformers Through Conditioned Embedded Tokens
by: Saratchandran, Hemanth, et al.
Published: (2025)
by: Saratchandran, Hemanth, et al.
Published: (2025)
A Sampling Theory Perspective on Activations for Implicit Neural Representations
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
D'OH: Decoder-Only Random Hypernetworks for Implicit Neural Representations
by: Gordon, Cameron, et al.
Published: (2024)
by: Gordon, Cameron, et al.
Published: (2024)
SineProject: Machine Unlearning for Stable Vision Language Alignment
by: Garg, Arpit, et al.
Published: (2025)
by: Garg, Arpit, et al.
Published: (2025)
Weight Conditioning for Smooth Optimization of Neural Networks
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
Transformers Pretrained on Procedural Data Contain Modular Structures for Algorithmic Reasoning
by: Shinnick, Zachary, et al.
Published: (2025)
by: Shinnick, Zachary, et al.
Published: (2025)
Mask the Target: A Plug-and-Play Regularizer Against LoRA Forgetting
by: Xu, Runze, et al.
Published: (2026)
by: Xu, Runze, et al.
Published: (2026)
Data Denoising and Derivative Estimation for Data-Driven Modeling of Nonlinear Dynamical Systems
by: Yao, Jiaqi, et al.
Published: (2025)
by: Yao, Jiaqi, et al.
Published: (2025)
Invertible Neural Warp for NeRF
by: Chng, Shin-Fang, et al.
Published: (2024)
by: Chng, Shin-Fang, et al.
Published: (2024)
Structured Initialization for Attention in Vision Transformers
by: Zheng, Jianqiao, et al.
Published: (2024)
by: Zheng, Jianqiao, et al.
Published: (2024)
Convolutional Initialization for Data-Efficient Vision Transformers
by: Zheng, Jianqiao, et al.
Published: (2024)
by: Zheng, Jianqiao, et al.
Published: (2024)
Procedural Pretraining: Warming Up Language Models with Abstract Data
by: Jiang, Liangze, et al.
Published: (2026)
by: Jiang, Liangze, et al.
Published: (2026)
Gradient Descent as a Shrinkage Operator for Spectral Bias
by: Lucey, Simon
Published: (2025)
by: Lucey, Simon
Published: (2025)
Spatioformer: A Geo-encoded Transformer for Large-Scale Plant Species Richness Prediction
by: Guo, Yiqing, et al.
Published: (2024)
by: Guo, Yiqing, et al.
Published: (2024)
Flashbacks to Harmonize Stability and Plasticity in Continual Learning
by: Mahmoodi, Leila, et al.
Published: (2025)
by: Mahmoodi, Leila, et al.
Published: (2025)
Towards Higher Effective Rank in Parameter-efficient Fine-tuning using Khatri--Rao Product
by: Albert, Paul, et al.
Published: (2025)
by: Albert, Paul, et al.
Published: (2025)
ReadMOF: Structure-Free Semantic Embeddings from Systematic MOF Nomenclature for Machine Learning
by: Zhu, Kewei, et al.
Published: (2026)
by: Zhu, Kewei, et al.
Published: (2026)
Inductive Graph Few-shot Class Incremental Learning
by: Li, Yayong, et al.
Published: (2024)
by: Li, Yayong, et al.
Published: (2024)
GPT Carry-On: Training Foundation Model for Customization Could Be Simple, Scalable and Affordable
by: Wangni, Jianqiao
Published: (2025)
by: Wangni, Jianqiao
Published: (2025)
Plant species richness prediction from DESIS hyperspectral data: A comparison study on feature extraction procedures and regression models
by: Guo, Yiqing, et al.
Published: (2023)
by: Guo, Yiqing, et al.
Published: (2023)
Enhancing Linear Attention with Residual Learning
by: Lai, Xunhao, et al.
Published: (2025)
by: Lai, Xunhao, et al.
Published: (2025)
Large-Scale In-Game Outcome Forecasting for Match, Team and Players in Football using an Axial Transformer Neural Network
by: Horton, Michael, et al.
Published: (2025)
by: Horton, Michael, et al.
Published: (2025)
Pre-training with Random Orthogonal Projection Image Modeling
by: Haghighat, Maryam, et al.
Published: (2023)
by: Haghighat, Maryam, et al.
Published: (2023)
Learning to Skip the Middle Layers of Transformers
by: Lawson, Tim, et al.
Published: (2025)
by: Lawson, Tim, et al.
Published: (2025)
Similar Items
-
Always Skip Attention
by: Ji, Yiping, et al.
Published: (2025) -
The Quantization Benefits of Residual-Free Transformers
by: Ji, Yiping, et al.
Published: (2026) -
Rethinking Attention: Polynomial Alternatives to Softmax in Transformers
by: Saratchandran, Hemanth, et al.
Published: (2024) -
The Inlet Rank Collapse in Implicit Neural Representations: Diagnosis and Unified Remedy
by: Zheng, Jianqiao, et al.
Published: (2026) -
From Tables to Signals: Revealing Spectral Adaptivity in TabPFN
by: Zheng, Jianqiao, et al.
Published: (2025)