The underlying structures of self-attention: symmetry, directionality, and emergent dynamics in Transformer training
Fuente:
arXiv
Guardado en:
| Autores principales: | Saponati, Matteo, Sager, Pascal, Aceituno, Pau Vilimelis, Stadelmann, Thilo, Grewe, Benjamin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A combination of noise and bilateral filters achieve supralinear and scalable adversarial robustness in CNNs
por: Stalder, Nicolas, et al.
Publicado: (2026)
por: Stalder, Nicolas, et al.
Publicado: (2026)
Continual Learning through Control Minimization
por: de Haan, Sander, et al.
Publicado: (2026)
por: de Haan, Sander, et al.
Publicado: (2026)
The Cooperative Network Architecture: Learning Structured Networks as Representation of Sensory Patterns
por: Sager, Pascal J., et al.
Publicado: (2024)
por: Sager, Pascal J., et al.
Publicado: (2024)
The Role of Temporal Hierarchy in Spiking Neural Networks
por: Moro, Filippo, et al.
Publicado: (2024)
por: Moro, Filippo, et al.
Publicado: (2024)
Deep Retrieval at CheckThat! 2025: Identifying Scientific Papers from Implicit Social Media Mentions via Hybrid Retrieval and Re-Ranking
por: Sager, Pascal J., et al.
Publicado: (2025)
por: Sager, Pascal J., et al.
Publicado: (2025)
Mixed-signal implementation of feedback-control optimizer for single-layer Spiking Neural Networks
por: Haag, Jonathan, et al.
Publicado: (2026)
por: Haag, Jonathan, et al.
Publicado: (2026)
Temporal horizons in forecasting: a performance-learnability trade-off
por: Aceituno, Pau Vilimelis, et al.
Publicado: (2025)
por: Aceituno, Pau Vilimelis, et al.
Publicado: (2025)
The emergence of clusters in self-attention dynamics
por: Geshkovski, Borjan, et al.
Publicado: (2023)
por: Geshkovski, Borjan, et al.
Publicado: (2023)
A feedback control optimizer for online and hardware-aware training of Spiking Neural Networks
por: Saponati, Matteo, et al.
Publicado: (2026)
por: Saponati, Matteo, et al.
Publicado: (2026)
Two types of pyramidal cells and their role in temporal processing
por: Vo, Anh Duong, et al.
Publicado: (2023)
por: Vo, Anh Duong, et al.
Publicado: (2023)
A Comprehensive Survey of Deep Transfer Learning for Anomaly Detection in Industrial Time Series: Methods, Applications, and Directions
por: Yan, Peng, et al.
Publicado: (2023)
por: Yan, Peng, et al.
Publicado: (2023)
Deep Neural Networks for Automatic Speaker Recognition Do Not Learn Supra-Segmental Temporal Features
por: Neururer, Daniel, et al.
Publicado: (2023)
por: Neururer, Daniel, et al.
Publicado: (2023)
A calibration framework to improve mechanistic forecasts with hybrid dynamic models
por: Victor Boussange, et al.
Publicado: (2025)
por: Victor Boussange, et al.
Publicado: (2025)
Learning Actionable World Models for Industrial Process Control
por: Yan, Peng, et al.
Publicado: (2025)
por: Yan, Peng, et al.
Publicado: (2025)
Multi-stage Bayesian optimisation for dynamic decision-making in self-driving labs
por: Torresi, Luca, et al.
Publicado: (2025)
por: Torresi, Luca, et al.
Publicado: (2025)
An extension of linear self-attention for in-context learning
por: Hagiwara, Katsuyuki
Publicado: (2025)
por: Hagiwara, Katsuyuki
Publicado: (2025)
Tucker Attention: A generalization of approximate attention mechanisms
por: Klein, Timon, et al.
Publicado: (2026)
por: Klein, Timon, et al.
Publicado: (2026)
Poly-attention: a general scheme for higher-order self-attention
por: Chakrabarti, Sayak, et al.
Publicado: (2026)
por: Chakrabarti, Sayak, et al.
Publicado: (2026)
Mechanistic Interpretability of RNNs emulating Hidden Markov Models
por: Torre, Elia, et al.
Publicado: (2025)
por: Torre, Elia, et al.
Publicado: (2025)
Testing the spin-bath view of self-attention: A Hamiltonian analysis of GPT-2 Transformer
por: Bhattacharjee, Satadeep, et al.
Publicado: (2025)
por: Bhattacharjee, Satadeep, et al.
Publicado: (2025)
Intrinsic Biologically Plausible Adversarial Robustness
por: Farinha, Matilde Tristany, et al.
Publicado: (2023)
por: Farinha, Matilde Tristany, et al.
Publicado: (2023)
A document is worth a structured record: Principled inductive bias design for document recognition
por: Meyer, Benjamin, et al.
Publicado: (2025)
por: Meyer, Benjamin, et al.
Publicado: (2025)
D3MES: Diffusion Transformer with multihead equivariant self-attention for 3D molecule generation
por: Zhang, Zhejun, et al.
Publicado: (2025)
por: Zhang, Zhejun, et al.
Publicado: (2025)
Local vs Global continual learning
por: Lanzillotta, Giulia, et al.
Publicado: (2024)
por: Lanzillotta, Giulia, et al.
Publicado: (2024)
Guided Transfer Learning for Discrete Diffusion Models
por: Kleutgens, Julian, et al.
Publicado: (2025)
por: Kleutgens, Julian, et al.
Publicado: (2025)
Learning spatially structured open quantum dynamics with regional-attention transformers
por: Du, Dounan, et al.
Publicado: (2025)
por: Du, Dounan, et al.
Publicado: (2025)
Dynamic metastability in the self-attention model
por: Geshkovski, Borjan, et al.
Publicado: (2024)
por: Geshkovski, Borjan, et al.
Publicado: (2024)
Myosotis: structured computation for attention like layer
por: Egorov, Evgenii, et al.
Publicado: (2025)
por: Egorov, Evgenii, et al.
Publicado: (2025)
Improving self-training under distribution shifts via anchored confidence with theoretical guarantees
por: Joo, Taejong, et al.
Publicado: (2024)
por: Joo, Taejong, et al.
Publicado: (2024)
The emergence of sparse attention: impact of data distribution and benefits of repetition
por: Zucchet, Nicolas, et al.
Publicado: (2025)
por: Zucchet, Nicolas, et al.
Publicado: (2025)
Probing self-attention in self-supervised speech models for cross-linguistic differences
por: Gopinath, Sai, et al.
Publicado: (2024)
por: Gopinath, Sai, et al.
Publicado: (2024)
Homomorphism Autoencoder -- Learning Group Structured Representations from Observed Transitions
por: Keurti, Hamza, et al.
Publicado: (2022)
por: Keurti, Hamza, et al.
Publicado: (2022)
A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions
por: Sager, Pascal J., et al.
Publicado: (2025)
por: Sager, Pascal J., et al.
Publicado: (2025)
Abstractors and relational cross-attention: An inductive bias for explicit relational reasoning in Transformers
por: Altabaa, Awni, et al.
Publicado: (2023)
por: Altabaa, Awni, et al.
Publicado: (2023)
Efficient training for large-scale optical neural network using an evolutionary strategy and attention pruning
por: Yang, Zhiwei, et al.
Publicado: (2025)
por: Yang, Zhiwei, et al.
Publicado: (2025)
Spatially-informed transformers: Injecting geostatistical covariance biases into self-attention for spatio-temporal forecasting
por: Calleo, Yuri
Publicado: (2025)
por: Calleo, Yuri
Publicado: (2025)
Self-attention as an attractor network: transient memories without backpropagation
por: D'Amico, Francesco, et al.
Publicado: (2024)
por: D'Amico, Francesco, et al.
Publicado: (2024)
TransformerFAM: Feedback attention is working memory
por: Hwang, Dongseong, et al.
Publicado: (2024)
por: Hwang, Dongseong, et al.
Publicado: (2024)
Universal and efficient graph neural networks with dynamic attention for machine learning interatomic potentials
por: Bi, Shuyu, et al.
Publicado: (2026)
por: Bi, Shuyu, et al.
Publicado: (2026)
Inversion dynamics of class manifolds in deep learning reveals tradeoffs underlying generalisation
por: Ciceri, Simone, et al.
Publicado: (2023)
por: Ciceri, Simone, et al.
Publicado: (2023)
Ejemplares similares
-
A combination of noise and bilateral filters achieve supralinear and scalable adversarial robustness in CNNs
por: Stalder, Nicolas, et al.
Publicado: (2026) -
Continual Learning through Control Minimization
por: de Haan, Sander, et al.
Publicado: (2026) -
The Cooperative Network Architecture: Learning Structured Networks as Representation of Sensory Patterns
por: Sager, Pascal J., et al.
Publicado: (2024) -
The Role of Temporal Hierarchy in Spiking Neural Networks
por: Moro, Filippo, et al.
Publicado: (2024) -
Deep Retrieval at CheckThat! 2025: Identifying Scientific Papers from Implicit Social Media Mentions via Hybrid Retrieval and Re-Ranking
por: Sager, Pascal J., et al.
Publicado: (2025)