Unification of Symmetries Inside Neural Networks: Transformer, Feedforward and Neural ODE

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hashimoto, Koji, Hirono, Yuji, Sannai, Akiyoshi
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913222999670784
author Hashimoto, Koji
Hirono, Yuji
Sannai, Akiyoshi
author_facet Hashimoto, Koji
Hirono, Yuji
Sannai, Akiyoshi
contents Understanding the inner workings of neural networks, including transformers, remains one of the most challenging puzzles in machine learning. This study introduces a novel approach by applying the principles of gauge symmetries, a key concept in physics, to neural network architectures. By regarding model functions as physical observables, we find that parametric redundancies of various machine learning models can be interpreted as gauge symmetries. We mathematically formulate the parametric redundancies in neural ODEs, and find that their gauge symmetries are given by spacetime diffeomorphisms, which play a fundamental role in Einstein's theory of gravity. Viewing neural ODEs as a continuum version of feedforward neural networks, we show that the parametric redundancies in feedforward neural networks are indeed lifted to diffeomorphisms in neural ODEs. We further extend our analysis to transformer models, finding natural correspondences with neural ODEs and their gauge symmetries. The concept of gauge symmetries sheds light on the complex behavior of deep learning models through physics and provides us with a unifying perspective for analyzing various machine learning architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2402_02362
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unification of Symmetries Inside Neural Networks: Transformer, Feedforward and Neural ODE
Hashimoto, Koji
Hirono, Yuji
Sannai, Akiyoshi
Machine Learning
Artificial Intelligence
High Energy Physics - Theory
Computational Physics
Understanding the inner workings of neural networks, including transformers, remains one of the most challenging puzzles in machine learning. This study introduces a novel approach by applying the principles of gauge symmetries, a key concept in physics, to neural network architectures. By regarding model functions as physical observables, we find that parametric redundancies of various machine learning models can be interpreted as gauge symmetries. We mathematically formulate the parametric redundancies in neural ODEs, and find that their gauge symmetries are given by spacetime diffeomorphisms, which play a fundamental role in Einstein's theory of gravity. Viewing neural ODEs as a continuum version of feedforward neural networks, we show that the parametric redundancies in feedforward neural networks are indeed lifted to diffeomorphisms in neural ODEs. We further extend our analysis to transformer models, finding natural correspondences with neural ODEs and their gauge symmetries. The concept of gauge symmetries sheds light on the complex behavior of deep learning models through physics and provides us with a unifying perspective for analyzing various machine learning architectures.
title Unification of Symmetries Inside Neural Networks: Transformer, Feedforward and Neural ODE
topic Machine Learning
Artificial Intelligence
High Energy Physics - Theory
Computational Physics
url https://arxiv.org/abs/2402.02362