Segment-Level Diffusion: A Framework for Controllable Long-Form Generation with Diffusion Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Xiaochen, Karadzhov, Georgi, Whitehouse, Chenxi, Vlachos, Andreas
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909621671690240
author Zhu, Xiaochen
Karadzhov, Georgi
Whitehouse, Chenxi
Vlachos, Andreas
author_facet Zhu, Xiaochen
Karadzhov, Georgi
Whitehouse, Chenxi
Vlachos, Andreas
contents Diffusion models have shown promise in text generation, but often struggle with generating long, coherent, and contextually accurate text. Token-level diffusion doesn't model word-order dependencies explicitly and operates on short, fixed output windows, while passage-level diffusion struggles with learning robust representations for long-form text. To address these challenges, we propose Segment-Level Diffusion (SLD), a framework that enhances diffusion-based text generation through text segmentation, robust representation training with adversarial and contrastive learning, and improved latent-space guidance. By segmenting long-form outputs into multiple latent representations and decoding them with an autoregressive decoder, SLD simplifies diffusion predictions and improves scalability. Experiments on four datasets demonstrate that, when compared to other diffusion and autoregressive baselines SLD achieves competitive or superior fluency, coherence, and contextual compatibility in automatic and human evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11333
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Segment-Level Diffusion: A Framework for Controllable Long-Form Generation with Diffusion Language Models
Zhu, Xiaochen
Karadzhov, Georgi
Whitehouse, Chenxi
Vlachos, Andreas
Computation and Language
Artificial Intelligence
Diffusion models have shown promise in text generation, but often struggle with generating long, coherent, and contextually accurate text. Token-level diffusion doesn't model word-order dependencies explicitly and operates on short, fixed output windows, while passage-level diffusion struggles with learning robust representations for long-form text. To address these challenges, we propose Segment-Level Diffusion (SLD), a framework that enhances diffusion-based text generation through text segmentation, robust representation training with adversarial and contrastive learning, and improved latent-space guidance. By segmenting long-form outputs into multiple latent representations and decoding them with an autoregressive decoder, SLD simplifies diffusion predictions and improves scalability. Experiments on four datasets demonstrate that, when compared to other diffusion and autoregressive baselines SLD achieves competitive or superior fluency, coherence, and contextual compatibility in automatic and human evaluations.
title Segment-Level Diffusion: A Framework for Controllable Long-Form Generation with Diffusion Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2412.11333