OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, John, Havasi, Marton, Berrada, Tariq, Zettlemoyer, Luke, Chen, Ricky T. Q.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912756641300480
author Nguyen, John
Havasi, Marton
Berrada, Tariq
Zettlemoyer, Luke
Chen, Ricky T. Q.
author_facet Nguyen, John
Havasi, Marton
Berrada, Tariq
Zettlemoyer, Luke
Chen, Ricky T. Q.
contents We present OneFlow, the first non-autoregressive multimodal model that enables variable-length and concurrent mixed-modal generation. Unlike autoregressive models that enforce rigid causal ordering between text and image generation, OneFlow combines an insertion-based Edit Flow for discrete text tokens with Flow Matching for image latents. OneFlow enables concurrent text-image synthesis with hierarchical sampling that prioritizes content over grammar. Through controlled experiments across model sizes from 1B to 8B, we demonstrate that OneFlow outperforms autoregressive baselines on both generation and understanding tasks while using up to 50% fewer training FLOPs. OneFlow surpasses both autoregressive and diffusion-based approaches while unlocking new capabilities for concurrent generation, iterative refinement, and natural reasoning-like generation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03506
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
Nguyen, John
Havasi, Marton
Berrada, Tariq
Zettlemoyer, Luke
Chen, Ricky T. Q.
Artificial Intelligence
We present OneFlow, the first non-autoregressive multimodal model that enables variable-length and concurrent mixed-modal generation. Unlike autoregressive models that enforce rigid causal ordering between text and image generation, OneFlow combines an insertion-based Edit Flow for discrete text tokens with Flow Matching for image latents. OneFlow enables concurrent text-image synthesis with hierarchical sampling that prioritizes content over grammar. Through controlled experiments across model sizes from 1B to 8B, we demonstrate that OneFlow outperforms autoregressive baselines on both generation and understanding tasks while using up to 50% fewer training FLOPs. OneFlow surpasses both autoregressive and diffusion-based approaches while unlocking new capabilities for concurrent generation, iterative refinement, and natural reasoning-like generation.
title OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
topic Artificial Intelligence
url https://arxiv.org/abs/2510.03506