Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive Flows

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Ruixiang, Zhai, Shuangfei, Gu, Jiatao, Zhang, Yizhe, Zheng, Huangjie, Chen, Tianrong, Bautista, Miguel Angel, Susskind, Josh, Jaitly, Navdeep
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908428875595776
author Zhang, Ruixiang
Zhai, Shuangfei
Gu, Jiatao
Zhang, Yizhe
Zheng, Huangjie
Chen, Tianrong
Bautista, Miguel Angel
Susskind, Josh
Jaitly, Navdeep
author_facet Zhang, Ruixiang
Zhai, Shuangfei
Gu, Jiatao
Zhang, Yizhe
Zheng, Huangjie
Chen, Tianrong
Bautista, Miguel Angel
Susskind, Josh
Jaitly, Navdeep
contents Autoregressive models have driven remarkable progress in language modeling. Their foundational reliance on discrete tokens, unidirectional context, and single-pass decoding, while central to their success, also inspires the exploration of a design space that could offer new axes of modeling flexibility. In this work, we explore an alternative paradigm, shifting language modeling from a discrete token space to a continuous latent space. We propose a novel framework TarFlowLM, that employs transformer-based autoregressive normalizing flows to model these continuous representations. This approach unlocks substantial flexibility, enabling the construction of models that can capture global bi-directional context through stacked, alternating-direction autoregressive transformations, support block-wise generation with flexible token patch sizes, and facilitate a hierarchical multi-pass generation process. We further propose new mixture-based coupling transformations designed to capture complex dependencies within the latent space shaped by discrete data, and demonstrate theoretical connections to conventional discrete autoregressive models. Extensive experiments on language modeling benchmarks demonstrate strong likelihood performance and highlight the flexible modeling capabilities inherent in our framework.
format Preprint
id arxiv_https___arxiv_org_abs_2507_00425
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive Flows
Zhang, Ruixiang
Zhai, Shuangfei
Gu, Jiatao
Zhang, Yizhe
Zheng, Huangjie
Chen, Tianrong
Bautista, Miguel Angel
Susskind, Josh
Jaitly, Navdeep
Machine Learning
Computation and Language
Autoregressive models have driven remarkable progress in language modeling. Their foundational reliance on discrete tokens, unidirectional context, and single-pass decoding, while central to their success, also inspires the exploration of a design space that could offer new axes of modeling flexibility. In this work, we explore an alternative paradigm, shifting language modeling from a discrete token space to a continuous latent space. We propose a novel framework TarFlowLM, that employs transformer-based autoregressive normalizing flows to model these continuous representations. This approach unlocks substantial flexibility, enabling the construction of models that can capture global bi-directional context through stacked, alternating-direction autoregressive transformations, support block-wise generation with flexible token patch sizes, and facilitate a hierarchical multi-pass generation process. We further propose new mixture-based coupling transformations designed to capture complex dependencies within the latent space shaped by discrete data, and demonstrate theoretical connections to conventional discrete autoregressive models. Extensive experiments on language modeling benchmarks demonstrate strong likelihood performance and highlight the flexible modeling capabilities inherent in our framework.
title Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive Flows
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2507.00425