Why "classic" Transformers are shallow and how to make them go deep
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Yueyao, Zhang, Yin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Why Flow Matching is Particle Swarm Optimization?
von: Ouyang, Kaichen
Veröffentlicht: (2025)
von: Ouyang, Kaichen
Veröffentlicht: (2025)
Sample-based Dynamic Hierarchical Transformer with Layer and Head Flexibility via Contextual Bandit
von: Meng, Fanfei, et al.
Veröffentlicht: (2023)
von: Meng, Fanfei, et al.
Veröffentlicht: (2023)
An algorithmic framework for the optimization of deep neural networks architectures and hyperparameters
von: Keisler, Julie, et al.
Veröffentlicht: (2023)
von: Keisler, Julie, et al.
Veröffentlicht: (2023)
Evolutionary fine tuning of quantized convolution-based deep learning models
von: Pietroń, Marcin
Veröffentlicht: (2026)
von: Pietroń, Marcin
Veröffentlicht: (2026)
SGHormer: An Energy-Saving Graph Transformer Driven by Spikes
von: Zhang, Huizhe, et al.
Veröffentlicht: (2024)
von: Zhang, Huizhe, et al.
Veröffentlicht: (2024)
From Frege to chatGPT: Compositionality in language, cognition, and deep neural networks
von: Russin, Jacob, et al.
Veröffentlicht: (2024)
von: Russin, Jacob, et al.
Veröffentlicht: (2024)
Spiking Point Transformer for Point Cloud Classification
von: Wu, Peixi, et al.
Veröffentlicht: (2025)
von: Wu, Peixi, et al.
Veröffentlicht: (2025)
Attending to Graph Transformers
von: Müller, Luis, et al.
Veröffentlicht: (2023)
von: Müller, Luis, et al.
Veröffentlicht: (2023)
Bayes-CATSI: A variational Bayesian deep learning framework for medical time series data imputation
von: Kulkarni, Omkar, et al.
Veröffentlicht: (2024)
von: Kulkarni, Omkar, et al.
Veröffentlicht: (2024)
Structure Development in List-Sorting Transformers
von: Urdshals, Einar, et al.
Veröffentlicht: (2025)
von: Urdshals, Einar, et al.
Veröffentlicht: (2025)
Investigating Recurrent Transformers with Dynamic Halt
von: Chowdhury, Jishnu Ray, et al.
Veröffentlicht: (2024)
von: Chowdhury, Jishnu Ray, et al.
Veröffentlicht: (2024)
Understanding Transformer Optimization via Gradient Heterogeneity
von: Tomihari, Akiyoshi, et al.
Veröffentlicht: (2025)
von: Tomihari, Akiyoshi, et al.
Veröffentlicht: (2025)
MoEUT: Mixture-of-Experts Universal Transformers
von: Csordás, Róbert, et al.
Veröffentlicht: (2024)
von: Csordás, Róbert, et al.
Veröffentlicht: (2024)
General-Purpose In-Context Learning by Meta-Learning Transformers
von: Kirsch, Louis, et al.
Veröffentlicht: (2022)
von: Kirsch, Louis, et al.
Veröffentlicht: (2022)
QSViT: A Methodology for Quantizing Spiking Vision Transformers
von: Putra, Rachmad Vidya Wicaksana, et al.
Veröffentlicht: (2025)
von: Putra, Rachmad Vidya Wicaksana, et al.
Veröffentlicht: (2025)
Decision SpikeFormer: Spike-Driven Transformer for Decision Making
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain
von: Kosowski, Adrian, et al.
Veröffentlicht: (2025)
von: Kosowski, Adrian, et al.
Veröffentlicht: (2025)
Weight Decay Regimes in Grokking Transformers: Cheap Online Diagnostics
von: Verma, Lucky
Veröffentlicht: (2026)
von: Verma, Lucky
Veröffentlicht: (2026)
Sparsity Moves Computation: How FFN Architecture Reshapes Attention in Small Transformers
von: Smithline, Gabriel, et al.
Veröffentlicht: (2026)
von: Smithline, Gabriel, et al.
Veröffentlicht: (2026)
Enabling Robust In-Context Memory and Rapid Task Adaptation in Transformers with Hebbian and Gradient-Based Plasticity
von: Chaudhary, Siddharth
Veröffentlicht: (2025)
von: Chaudhary, Siddharth
Veröffentlicht: (2025)
Enhancing Graph Representation Learning with Attention-Driven Spiking Neural Networks
von: Yin, Huifeng, et al.
Veröffentlicht: (2024)
von: Yin, Huifeng, et al.
Veröffentlicht: (2024)
Decoding Listeners Identity: Person Identification from EEG Signals Using a Lightweight Spiking Transformer
von: Lin, Zheyuan, et al.
Veröffentlicht: (2025)
von: Lin, Zheyuan, et al.
Veröffentlicht: (2025)
Bridging Synthetic and Real Routing Problems via LLM-Guided Instance Generation and Progressive Adaptation
von: Zhu, Jianghan, et al.
Veröffentlicht: (2025)
von: Zhu, Jianghan, et al.
Veröffentlicht: (2025)
Why Fine-Tuning Encourages Hallucinations and How to Fix It
von: Kaplan, Guy, et al.
Veröffentlicht: (2026)
von: Kaplan, Guy, et al.
Veröffentlicht: (2026)
Predicting Deterioration in Mild Cognitive Impairment with Survival Transformers, Extreme Gradient Boosting and Cox Proportional Hazard Modelling
von: Musto, Henry, et al.
Veröffentlicht: (2024)
von: Musto, Henry, et al.
Veröffentlicht: (2024)
Provably Optimal Memory Capacity for Modern Hopfield Models: Transformer-Compatible Dense Associative Memories as Spherical Codes
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2024)
CogDPM: Diffusion Probabilistic Models via Cognitive Predictive Coding
von: Chen, Kaiyuan, et al.
Veröffentlicht: (2024)
von: Chen, Kaiyuan, et al.
Veröffentlicht: (2024)
SkipSNN: Efficiently Classifying Spike Trains with Event-attention
von: Yin, Hang, et al.
Veröffentlicht: (2024)
von: Yin, Hang, et al.
Veröffentlicht: (2024)
STARS: Spike Tail-Aware Relational Synthesis for ANN-to-SNN Data-Free Knowledge Distillation
von: Ye, Shuhan, et al.
Veröffentlicht: (2026)
von: Ye, Shuhan, et al.
Veröffentlicht: (2026)
Dynamic Spiking Framework for Graph Neural Networks
von: Yin, Nan, et al.
Veröffentlicht: (2023)
von: Yin, Nan, et al.
Veröffentlicht: (2023)
Seemingly Redundant Modules Enhance Robust Odor Learning in Fruit Flies
von: Li, Haiyang, et al.
Veröffentlicht: (2025)
von: Li, Haiyang, et al.
Veröffentlicht: (2025)
Continuous Spiking Graph Neural Networks
von: Yin, Nan, et al.
Veröffentlicht: (2024)
von: Yin, Nan, et al.
Veröffentlicht: (2024)
Evolutionary Ensemble of Agents
von: Yu, Zongmin, et al.
Veröffentlicht: (2026)
von: Yu, Zongmin, et al.
Veröffentlicht: (2026)
Unveiling the Potential of Spiking Dynamics in Graph Representation Learning through Spatial-Temporal Normalization and Coding Strategies
von: Xu, Mingkun, et al.
Veröffentlicht: (2024)
von: Xu, Mingkun, et al.
Veröffentlicht: (2024)
Language Models Learn Universal Representations of Numbers and Here's Why You Should Care
von: Štefánik, Michal, et al.
Veröffentlicht: (2025)
von: Štefánik, Michal, et al.
Veröffentlicht: (2025)
Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
Understanding the Functional Roles of Modelling Components in Spiking Neural Networks
von: Yin, Huifeng, et al.
Veröffentlicht: (2024)
von: Yin, Huifeng, et al.
Veröffentlicht: (2024)
Evolution Meets Diffusion: Efficient Neural Architecture Generation
von: Zhou, Bingye, et al.
Veröffentlicht: (2025)
von: Zhou, Bingye, et al.
Veröffentlicht: (2025)
On the Temperature of Machine Learning Systems
von: Zhang, Dong
Veröffentlicht: (2024)
von: Zhang, Dong
Veröffentlicht: (2024)
Advancing Direct Training for Spiking Neural Networks with Circulate-Firing Neurons and Learnable Gradients
von: Zhou, Feifan, et al.
Veröffentlicht: (2026)
von: Zhou, Feifan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Why Flow Matching is Particle Swarm Optimization?
von: Ouyang, Kaichen
Veröffentlicht: (2025) -
Sample-based Dynamic Hierarchical Transformer with Layer and Head Flexibility via Contextual Bandit
von: Meng, Fanfei, et al.
Veröffentlicht: (2023) -
An algorithmic framework for the optimization of deep neural networks architectures and hyperparameters
von: Keisler, Julie, et al.
Veröffentlicht: (2023) -
Evolutionary fine tuning of quantized convolution-based deep learning models
von: Pietroń, Marcin
Veröffentlicht: (2026) -
SGHormer: An Energy-Saving Graph Transformer Driven by Spikes
von: Zhang, Huizhe, et al.
Veröffentlicht: (2024)