Bridging the Empirical-Theoretical Gap in Neural Network Formal Language Learning Using Minimum Description Length

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lan, Nur, Chemla, Emmanuel, Katzir, Roni
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909217954201600
author Lan, Nur
Chemla, Emmanuel
Katzir, Roni
author_facet Lan, Nur
Chemla, Emmanuel
Katzir, Roni
contents Neural networks offer good approximation to many tasks but consistently fail to reach perfect generalization, even when theoretical work shows that such perfect solutions can be expressed by certain architectures. Using the task of formal language learning, we focus on one simple formal language and show that the theoretically correct solution is in fact not an optimum of commonly used objectives -- even with regularization techniques that according to common wisdom should lead to simple weights and good generalization (L1, L2) or other meta-heuristics (early-stopping, dropout). On the other hand, replacing standard targets with the Minimum Description Length objective (MDL) results in the correct solution being an optimum.
format Preprint
id arxiv_https___arxiv_org_abs_2402_10013
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Bridging the Empirical-Theoretical Gap in Neural Network Formal Language Learning Using Minimum Description Length
Lan, Nur
Chemla, Emmanuel
Katzir, Roni
Computation and Language
Formal Languages and Automata Theory
Neural networks offer good approximation to many tasks but consistently fail to reach perfect generalization, even when theoretical work shows that such perfect solutions can be expressed by certain architectures. Using the task of formal language learning, we focus on one simple formal language and show that the theoretically correct solution is in fact not an optimum of commonly used objectives -- even with regularization techniques that according to common wisdom should lead to simple weights and good generalization (L1, L2) or other meta-heuristics (early-stopping, dropout). On the other hand, replacing standard targets with the Minimum Description Length objective (MDL) results in the correct solution being an optimum.
title Bridging the Empirical-Theoretical Gap in Neural Network Formal Language Learning Using Minimum Description Length
topic Computation and Language
Formal Languages and Automata Theory
url https://arxiv.org/abs/2402.10013