Structure Development in List-Sorting Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Urdshals, Einar, Urdshals, Jasmina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Attending to Graph Transformers
by: Müller, Luis, et al.
Published: (2023)
by: Müller, Luis, et al.
Published: (2023)
Investigating Recurrent Transformers with Dynamic Halt
by: Chowdhury, Jishnu Ray, et al.
Published: (2024)
by: Chowdhury, Jishnu Ray, et al.
Published: (2024)
Understanding Transformer Optimization via Gradient Heterogeneity
by: Tomihari, Akiyoshi, et al.
Published: (2025)
by: Tomihari, Akiyoshi, et al.
Published: (2025)
Spiking Point Transformer for Point Cloud Classification
by: Wu, Peixi, et al.
Published: (2025)
by: Wu, Peixi, et al.
Published: (2025)
MoEUT: Mixture-of-Experts Universal Transformers
by: Csordás, Róbert, et al.
Published: (2024)
by: Csordás, Róbert, et al.
Published: (2024)
QSViT: A Methodology for Quantizing Spiking Vision Transformers
by: Putra, Rachmad Vidya Wicaksana, et al.
Published: (2025)
by: Putra, Rachmad Vidya Wicaksana, et al.
Published: (2025)
General-Purpose In-Context Learning by Meta-Learning Transformers
by: Kirsch, Louis, et al.
Published: (2022)
by: Kirsch, Louis, et al.
Published: (2022)
SGHormer: An Energy-Saving Graph Transformer Driven by Spikes
by: Zhang, Huizhe, et al.
Published: (2024)
by: Zhang, Huizhe, et al.
Published: (2024)
Decision SpikeFormer: Spike-Driven Transformer for Decision Making
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain
by: Kosowski, Adrian, et al.
Published: (2025)
by: Kosowski, Adrian, et al.
Published: (2025)
Weight Decay Regimes in Grokking Transformers: Cheap Online Diagnostics
by: Verma, Lucky
Published: (2026)
by: Verma, Lucky
Published: (2026)
Why "classic" Transformers are shallow and how to make them go deep
by: Yu, Yueyao, et al.
Published: (2023)
by: Yu, Yueyao, et al.
Published: (2023)
Sparsity Moves Computation: How FFN Architecture Reshapes Attention in Small Transformers
by: Smithline, Gabriel, et al.
Published: (2026)
by: Smithline, Gabriel, et al.
Published: (2026)
Structure of Artificial Neural Networks -- Empirical Investigations
by: Stier, Julian
Published: (2024)
by: Stier, Julian
Published: (2024)
Sample-based Dynamic Hierarchical Transformer with Layer and Head Flexibility via Contextual Bandit
by: Meng, Fanfei, et al.
Published: (2023)
by: Meng, Fanfei, et al.
Published: (2023)
Multi-Objective Quality-Diversity for Crystal Structure Prediction
by: Janmohamed, Hannah, et al.
Published: (2024)
by: Janmohamed, Hannah, et al.
Published: (2024)
Structure as Computation: Developmental Generation of Minimal Neural Circuits
by: Zhou, Duan
Published: (2026)
by: Zhou, Duan
Published: (2026)
Enabling Robust In-Context Memory and Rapid Task Adaptation in Transformers with Hebbian and Gradient-Based Plasticity
by: Chaudhary, Siddharth
Published: (2025)
by: Chaudhary, Siddharth
Published: (2025)
Structural Equation-VAE: Disentangled Latent Representations for Tabular Data
by: Zhang, Ruiyu, et al.
Published: (2025)
by: Zhang, Ruiyu, et al.
Published: (2025)
Provable Benefits of Complex Parameterizations for Structured State Space Models
by: Ran-Milo, Yuval, et al.
Published: (2024)
by: Ran-Milo, Yuval, et al.
Published: (2024)
Decoding Listeners Identity: Person Identification from EEG Signals Using a Lightweight Spiking Transformer
by: Lin, Zheyuan, et al.
Published: (2025)
by: Lin, Zheyuan, et al.
Published: (2025)
Scalable Learning in Structured Recurrent Spiking Neural Networks without Backpropagation
by: Tang, Bo, et al.
Published: (2026)
by: Tang, Bo, et al.
Published: (2026)
Structurally Flexible Neural Networks: Evolving the Building Blocks for General Agents
by: Pedersen, Joachim Winther, et al.
Published: (2024)
by: Pedersen, Joachim Winther, et al.
Published: (2024)
Predicting Deterioration in Mild Cognitive Impairment with Survival Transformers, Extreme Gradient Boosting and Cox Proportional Hazard Modelling
by: Musto, Henry, et al.
Published: (2024)
by: Musto, Henry, et al.
Published: (2024)
Provably Optimal Memory Capacity for Modern Hopfield Models: Transformer-Compatible Dense Associative Memories as Spherical Codes
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
SPEAR: Structured Pruning for Spiking Neural Networks via Synaptic Operation Estimation and Reinforcement Learning
by: Xie, Hui, et al.
Published: (2025)
by: Xie, Hui, et al.
Published: (2025)
A Deep Dive into Effects of Structural Bias on CMA-ES Performance along Affine Trajectories
by: van Stein, Niki, et al.
Published: (2024)
by: van Stein, Niki, et al.
Published: (2024)
SAN: Hypothesizing Long-Term Synaptic Development and Neural Engram Mechanism in Scalable Model's Parameter-Efficient Fine-Tuning
by: Dai, Gaole, et al.
Published: (2024)
by: Dai, Gaole, et al.
Published: (2024)
Statistical Properties of the King Wen Sequence: An Anti-Habituation Structure That Does Not Improve Neural Network Training
by: Chan, Augustin
Published: (2026)
by: Chan, Augustin
Published: (2026)
When Does Structure Matter in Continual Learning? Dimensionality Controls When Modularity Shapes Representational Geometry
by: Korte, Kathrin, et al.
Published: (2026)
by: Korte, Kathrin, et al.
Published: (2026)
Neuro-Symbolic Activation Discovery: Transferring Mathematical Structures from Physics to Ecology for Parameter-Efficient Neural Networks
by: Hajbi, Anas
Published: (2026)
by: Hajbi, Anas
Published: (2026)
Evolutionary Extreme Learning Machine of ab-initio Energy Landscapes for Crystal Structure Prediction using Manta Ray Optimization with Levy Flight
by: Rubio-Solis, Adrian
Published: (2026)
by: Rubio-Solis, Adrian
Published: (2026)
Sensitivity-Aware Mixed-Precision Quantization and Width Optimization of Deep Neural Networks Through Cluster-Based Tree-Structured Parzen Estimation
by: Azizi, Seyedarmin, et al.
Published: (2023)
by: Azizi, Seyedarmin, et al.
Published: (2023)
Datum-wise Transformer for Synthetic Tabular Data Detection in the Wild
by: Kindji, G. Charbel N., et al.
Published: (2025)
by: Kindji, G. Charbel N., et al.
Published: (2025)
GLU Attention Improve Transformer
by: Wang, Zehao
Published: (2025)
by: Wang, Zehao
Published: (2025)
Embodied Neuromorphic Artificial Intelligence for Robotics: Perspectives, Challenges, and Research Development Stack
by: Putra, Rachmad Vidya Wicaksana, et al.
Published: (2024)
by: Putra, Rachmad Vidya Wicaksana, et al.
Published: (2024)
A Transformer-based Neural Architecture Search Method
by: Wang, Shang, et al.
Published: (2025)
by: Wang, Shang, et al.
Published: (2025)
NOBLE: Accelerating Transformers with Nonlinear Low-Rank Branches
by: Smith, Ethan
Published: (2026)
by: Smith, Ethan
Published: (2026)
SNN4Agents: A Framework for Developing Energy-Efficient Embodied Spiking Neural Networks for Autonomous Agents
by: Putra, Rachmad Vidya Wicaksana, et al.
Published: (2024)
by: Putra, Rachmad Vidya Wicaksana, et al.
Published: (2024)
Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
by: Munkhdalai, Tsendsuren, et al.
Published: (2024)
by: Munkhdalai, Tsendsuren, et al.
Published: (2024)
Similar Items
-
Attending to Graph Transformers
by: Müller, Luis, et al.
Published: (2023) -
Investigating Recurrent Transformers with Dynamic Halt
by: Chowdhury, Jishnu Ray, et al.
Published: (2024) -
Understanding Transformer Optimization via Gradient Heterogeneity
by: Tomihari, Akiyoshi, et al.
Published: (2025) -
Spiking Point Transformer for Point Cloud Classification
by: Wu, Peixi, et al.
Published: (2025) -
MoEUT: Mixture-of-Experts Universal Transformers
by: Csordás, Róbert, et al.
Published: (2024)