Understanding Transformer Optimization via Gradient Heterogeneity
Fuente:
arXiv
Saved in:
| Main Authors: | Tomihari, Akiyoshi, Sato, Issei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from Jacobians
by: Tomihari, Akiyoshi, et al.
Published: (2025)
by: Tomihari, Akiyoshi, et al.
Published: (2025)
AlphaGrad: Non-Linear Gradient Normalization Optimizer
by: Sane, Soham
Published: (2025)
by: Sane, Soham
Published: (2025)
Enabling Robust In-Context Memory and Rapid Task Adaptation in Transformers with Hebbian and Gradient-Based Plasticity
by: Chaudhary, Siddharth
Published: (2025)
by: Chaudhary, Siddharth
Published: (2025)
A Learning-Based Cooperative Coevolution Framework for Heterogeneous Large-Scale Global Optimization
by: Qiu, Wenjie, et al.
Published: (2026)
by: Qiu, Wenjie, et al.
Published: (2026)
Predicting Deterioration in Mild Cognitive Impairment with Survival Transformers, Extreme Gradient Boosting and Cox Proportional Hazard Modelling
by: Musto, Henry, et al.
Published: (2024)
by: Musto, Henry, et al.
Published: (2024)
Understanding Deep Learning via Notions of Rank
by: Razin, Noam
Published: (2024)
by: Razin, Noam
Published: (2024)
Gradient-Free Training of Spiking Neural Networks via Low-Rank Evolution Strategies
by: Patankar, Dhruv, et al.
Published: (2026)
by: Patankar, Dhruv, et al.
Published: (2026)
Gradient-Free Continual Learning in Spiking Neural Networks via Inter-Spike Interval Regularization
by: Roy, Samrendra, et al.
Published: (2026)
by: Roy, Samrendra, et al.
Published: (2026)
Sample-based Dynamic Hierarchical Transformer with Layer and Head Flexibility via Contextual Bandit
by: Meng, Fanfei, et al.
Published: (2023)
by: Meng, Fanfei, et al.
Published: (2023)
Beyond Uniform Scaling: Exploring Depth Heterogeneity in Neural Architectures
by: T, Akash Guna R., et al.
Published: (2024)
by: T, Akash Guna R., et al.
Published: (2024)
A Truly Sparse and General Implementation of Gradient-Based Synaptic Plasticity
by: Lohoff, Jamie, et al.
Published: (2025)
by: Lohoff, Jamie, et al.
Published: (2025)
Towards Universal Offline Black-Box Optimization via Learning Language Model Embeddings
by: Tan, Rong-Xi, et al.
Published: (2025)
by: Tan, Rong-Xi, et al.
Published: (2025)
When Fireflies Cluster; Enhancing Automatic Clustering via Centroid-Guided Firefly Optimization
by: Ariyaratne, MKA, et al.
Published: (2026)
by: Ariyaratne, MKA, et al.
Published: (2026)
Robust Lagrangian and Adversarial Policy Gradient for Robust Constrained Markov Decision Processes
by: Bossens, David M.
Published: (2023)
by: Bossens, David M.
Published: (2023)
Attending to Graph Transformers
by: Müller, Luis, et al.
Published: (2023)
by: Müller, Luis, et al.
Published: (2023)
Advancing Direct Training for Spiking Neural Networks with Circulate-Firing Neurons and Learnable Gradients
by: Zhou, Feifan, et al.
Published: (2026)
by: Zhou, Feifan, et al.
Published: (2026)
Rethinking LLM-Driven Heuristic Design: Generating Efficient and Specialized Solvers via Dynamics-Aware Optimization
by: Wang, Rongzheng, et al.
Published: (2026)
by: Wang, Rongzheng, et al.
Published: (2026)
Scaling Policy Gradient Quality-Diversity with Massive Parallelization via Behavioral Variations
by: Mitsides, Konstantinos, et al.
Published: (2025)
by: Mitsides, Konstantinos, et al.
Published: (2025)
Understanding the Functional Roles of Modelling Components in Spiking Neural Networks
by: Yin, Huifeng, et al.
Published: (2024)
by: Yin, Huifeng, et al.
Published: (2024)
Structure Development in List-Sorting Transformers
by: Urdshals, Einar, et al.
Published: (2025)
by: Urdshals, Einar, et al.
Published: (2025)
Investigating Recurrent Transformers with Dynamic Halt
by: Chowdhury, Jishnu Ray, et al.
Published: (2024)
by: Chowdhury, Jishnu Ray, et al.
Published: (2024)
Evolutionary Computation and Explainable AI: A Roadmap to Understandable Intelligent Systems
by: Zhou, Ryan, et al.
Published: (2024)
by: Zhou, Ryan, et al.
Published: (2024)
Spiking Point Transformer for Point Cloud Classification
by: Wu, Peixi, et al.
Published: (2025)
by: Wu, Peixi, et al.
Published: (2025)
MoEUT: Mixture-of-Experts Universal Transformers
by: Csordás, Róbert, et al.
Published: (2024)
by: Csordás, Róbert, et al.
Published: (2024)
QSViT: A Methodology for Quantizing Spiking Vision Transformers
by: Putra, Rachmad Vidya Wicaksana, et al.
Published: (2025)
by: Putra, Rachmad Vidya Wicaksana, et al.
Published: (2025)
General-Purpose In-Context Learning by Meta-Learning Transformers
by: Kirsch, Louis, et al.
Published: (2022)
by: Kirsch, Louis, et al.
Published: (2022)
SGHormer: An Energy-Saving Graph Transformer Driven by Spikes
by: Zhang, Huizhe, et al.
Published: (2024)
by: Zhang, Huizhe, et al.
Published: (2024)
Multi-Timescale Conductance Spiking Networks: A Sparse, Gradient-Trainable Framework with Rich Firing Dynamics for Enhanced Temporal Processing
by: Fulleda-Garcia, Alex, et al.
Published: (2026)
by: Fulleda-Garcia, Alex, et al.
Published: (2026)
Offline Multi-Objective Optimization
by: Xue, Ke, et al.
Published: (2024)
by: Xue, Ke, et al.
Published: (2024)
Decision SpikeFormer: Spike-Driven Transformer for Decision Making
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain
by: Kosowski, Adrian, et al.
Published: (2025)
by: Kosowski, Adrian, et al.
Published: (2025)
Weight Decay Regimes in Grokking Transformers: Cheap Online Diagnostics
by: Verma, Lucky
Published: (2026)
by: Verma, Lucky
Published: (2026)
Surrogate Benchmarks for Model Merging Optimization
by: Akizuki, Rio, et al.
Published: (2025)
by: Akizuki, Rio, et al.
Published: (2025)
Few-shot Quality-Diversity Optimization
by: Salehi, Achkan, et al.
Published: (2021)
by: Salehi, Achkan, et al.
Published: (2021)
Reinforced In-Context Black-Box Optimization
by: Song, Lei, et al.
Published: (2024)
by: Song, Lei, et al.
Published: (2024)
Fast Fourier Transform-Based Spectral and Temporal Gradient Filtering for Differential Privacy
by: Shin, Hyeju, et al.
Published: (2025)
by: Shin, Hyeju, et al.
Published: (2025)
Why "classic" Transformers are shallow and how to make them go deep
by: Yu, Yueyao, et al.
Published: (2023)
by: Yu, Yueyao, et al.
Published: (2023)
Automated Algorithm Design for Auto-Tuning Optimizers
by: Willemsen, Floris-Jan, et al.
Published: (2025)
by: Willemsen, Floris-Jan, et al.
Published: (2025)
Why Flow Matching is Particle Swarm Optimization?
by: Ouyang, Kaichen
Published: (2025)
by: Ouyang, Kaichen
Published: (2025)
Multi-Task Optimization over Networks of Tasks
by: Hatzky, Julian, et al.
Published: (2026)
by: Hatzky, Julian, et al.
Published: (2026)
Similar Items
-
Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from Jacobians
by: Tomihari, Akiyoshi, et al.
Published: (2025) -
AlphaGrad: Non-Linear Gradient Normalization Optimizer
by: Sane, Soham
Published: (2025) -
Enabling Robust In-Context Memory and Rapid Task Adaptation in Transformers with Hebbian and Gradient-Based Plasticity
by: Chaudhary, Siddharth
Published: (2025) -
A Learning-Based Cooperative Coevolution Framework for Heterogeneous Large-Scale Global Optimization
by: Qiu, Wenjie, et al.
Published: (2026) -
Predicting Deterioration in Mild Cognitive Impairment with Survival Transformers, Extreme Gradient Boosting and Cox Proportional Hazard Modelling
by: Musto, Henry, et al.
Published: (2024)