Length Generalization with Log-Depth Recurrent Units

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Pert, Charles, Alrajeh, Dalal, Russo, Alessandra
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916045331103744
author Pert, Charles
Alrajeh, Dalal
Russo, Alessandra
author_facet Pert, Charles
Alrajeh, Dalal
Russo, Alessandra
contents Length generalization remains a persistent challenge for neural networks: recurrent models tend to suffer from positional biases, while transformers are constrained by fixed computational depth. Regular languages provide a frequently used testbed for evaluating length generalization, as label prediction can be checked for any sequence length. We propose MLP-LDRU, a type of Log-Depth Recurrent Unit, which captures a class of associativity-biased operators designed to approximate recurrence through parallel reduction. We evaluate MLP-LDRU on 21 regular-language tasks, consisting of standard benchmarks and new prefix languages, where it achieves 100% out-of-distribution accuracy on 18 tasks and at least 99.9% on the remaining 3 when increasing max training length, outperforming comparable recurrent and attention-based models. We further evaluate MLP-LDRU beyond regular languages on ListOps and NLP classification benchmarks, where it performs competitively.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26035
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Length Generalization with Log-Depth Recurrent Units
Pert, Charles
Alrajeh, Dalal
Russo, Alessandra
Machine Learning
Length generalization remains a persistent challenge for neural networks: recurrent models tend to suffer from positional biases, while transformers are constrained by fixed computational depth. Regular languages provide a frequently used testbed for evaluating length generalization, as label prediction can be checked for any sequence length. We propose MLP-LDRU, a type of Log-Depth Recurrent Unit, which captures a class of associativity-biased operators designed to approximate recurrence through parallel reduction. We evaluate MLP-LDRU on 21 regular-language tasks, consisting of standard benchmarks and new prefix languages, where it achieves 100% out-of-distribution accuracy on 18 tasks and at least 99.9% on the remaining 3 when increasing max training length, outperforming comparable recurrent and attention-based models. We further evaluate MLP-LDRU beyond regular languages on ListOps and NLP classification benchmarks, where it performs competitively.
title Length Generalization with Log-Depth Recurrent Units
topic Machine Learning
url https://arxiv.org/abs/2605.26035