Investigating Recurrent Transformers with Dynamic Halt

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chowdhury, Jishnu Ray, Caragea, Cornelia
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915111330906112
author Chowdhury, Jishnu Ray
Caragea, Cornelia
author_facet Chowdhury, Jishnu Ray
Caragea, Cornelia
contents In this paper, we comprehensively study the inductive biases of two major approaches to augmenting Transformers with a recurrent mechanism: (1) the approach of incorporating a depth-wise recurrence similar to Universal Transformers; and (2) the approach of incorporating a chunk-wise temporal recurrence like Temporal Latent Bottleneck. Furthermore, we propose and investigate novel ways to extend and combine the above methods - for example, we propose a global mean-based dynamic halting mechanism for Universal Transformers and an augmentation of Temporal Latent Bottleneck with elements from Universal Transformer. We compare the models and probe their inductive biases in several diagnostic tasks, such as Long Range Arena (LRA), flip-flop language modeling, ListOps, and Logical Inference. The code is released in: https://github.com/JRC1995/InvestigatingRecurrentTransformers/tree/main
format Preprint
id arxiv_https___arxiv_org_abs_2402_00976
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Investigating Recurrent Transformers with Dynamic Halt
Chowdhury, Jishnu Ray
Caragea, Cornelia
Machine Learning
Artificial Intelligence
Neural and Evolutionary Computing
In this paper, we comprehensively study the inductive biases of two major approaches to augmenting Transformers with a recurrent mechanism: (1) the approach of incorporating a depth-wise recurrence similar to Universal Transformers; and (2) the approach of incorporating a chunk-wise temporal recurrence like Temporal Latent Bottleneck. Furthermore, we propose and investigate novel ways to extend and combine the above methods - for example, we propose a global mean-based dynamic halting mechanism for Universal Transformers and an augmentation of Temporal Latent Bottleneck with elements from Universal Transformer. We compare the models and probe their inductive biases in several diagnostic tasks, such as Long Range Arena (LRA), flip-flop language modeling, ListOps, and Logical Inference. The code is released in: https://github.com/JRC1995/InvestigatingRecurrentTransformers/tree/main
title Investigating Recurrent Transformers with Dynamic Halt
topic Machine Learning
Artificial Intelligence
Neural and Evolutionary Computing
url https://arxiv.org/abs/2402.00976