AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Aleksandrov, Preslav, Kurmanji, Meghdad, Redondo, Fernando Garcia, O'Shea, David, Shen, William, Iacob, Alex, Sani, Lorenzo, Qiu, Xinchi, Cancedda, Nicola, Lane, Nicholas D.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909726215766016
author Aleksandrov, Preslav
Kurmanji, Meghdad
Redondo, Fernando Garcia
O'Shea, David
Shen, William
Iacob, Alex
Sani, Lorenzo
Qiu, Xinchi
Cancedda, Nicola
Lane, Nicholas D.
author_facet Aleksandrov, Preslav
Kurmanji, Meghdad
Redondo, Fernando Garcia
O'Shea, David
Shen, William
Iacob, Alex
Sani, Lorenzo
Qiu, Xinchi
Cancedda, Nicola
Lane, Nicholas D.
contents We introduce the Autoregressive Block-Based Iterative Encoder (AbbIE), a novel recursive generalization of the encoder-only Transformer architecture, which achieves better perplexity than a standard Transformer and allows for the dynamic scaling of compute resources at test time. This simple, recursive approach is a complement to scaling large language model (LLM) performance through parameter and token counts. AbbIE performs its iterations in latent space, but unlike latent reasoning models, does not require a specialized dataset or training protocol. We show that AbbIE upward generalizes (ability to generalize to arbitrary iteration lengths) at test time by only using 2 iterations during train time, far outperforming alternative iterative methods. AbbIE's ability to scale its computational expenditure based on the complexity of the task gives it an up to \textbf{12\%} improvement in zero-shot in-context learning tasks versus other iterative and standard methods and up to 5\% improvement in language perplexity. The results from this study open a new avenue to Transformer performance scaling. We perform all of our evaluations on model sizes up to 350M parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2507_08567
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling
Aleksandrov, Preslav
Kurmanji, Meghdad
Redondo, Fernando Garcia
O'Shea, David
Shen, William
Iacob, Alex
Sani, Lorenzo
Qiu, Xinchi
Cancedda, Nicola
Lane, Nicholas D.
Machine Learning
We introduce the Autoregressive Block-Based Iterative Encoder (AbbIE), a novel recursive generalization of the encoder-only Transformer architecture, which achieves better perplexity than a standard Transformer and allows for the dynamic scaling of compute resources at test time. This simple, recursive approach is a complement to scaling large language model (LLM) performance through parameter and token counts. AbbIE performs its iterations in latent space, but unlike latent reasoning models, does not require a specialized dataset or training protocol. We show that AbbIE upward generalizes (ability to generalize to arbitrary iteration lengths) at test time by only using 2 iterations during train time, far outperforming alternative iterative methods. AbbIE's ability to scale its computational expenditure based on the complexity of the task gives it an up to \textbf{12\%} improvement in zero-shot in-context learning tasks versus other iterative and standard methods and up to 5\% improvement in language perplexity. The results from this study open a new avenue to Transformer performance scaling. We perform all of our evaluations on model sizes up to 350M parameters.
title AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling
topic Machine Learning
url https://arxiv.org/abs/2507.08567