Diffusion Is Your Friend in Show, Suggest and Tell

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Jia Cheng, Cavicchioli, Roberto, Capotondi, Alessandro
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915667896172544
author Hu, Jia Cheng
Cavicchioli, Roberto
Capotondi, Alessandro
author_facet Hu, Jia Cheng
Cavicchioli, Roberto
Capotondi, Alessandro
contents Diffusion Denoising models demonstrated impressive results across generative Computer Vision tasks, but they still fail to outperform standard autoregressive solutions in the discrete domain, and only match them at best. In this work, we propose a different paradigm by adopting diffusion models to provide suggestions to the autoregressive generation rather than replacing them. By doing so, we combine the bidirectional and refining capabilities of the former with the strong linguistic structure provided by the latter. To showcase its effectiveness, we present Show, Suggest and Tell (SST), which achieves State-of-the-Art results on COCO, among models in a similar setting. In particular, SST achieves 125.1 CIDEr-D on the COCO dataset without Reinforcement Learning, outperforming both autoregressive and diffusion model State-of-the-Art results by 1.5 and 2.5 points. On top of the strong results, we performed extensive experiments to validate the proposal and analyze the impact of the suggestion module. Results demonstrate a positive correlation between suggestion and caption quality, overall indicating a currently underexplored but promising research direction. Code will be available at: https://github.com/jchenghu/show\_suggest\_tell.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10038
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Diffusion Is Your Friend in Show, Suggest and Tell
Hu, Jia Cheng
Cavicchioli, Roberto
Capotondi, Alessandro
Computer Vision and Pattern Recognition
Computation and Language
Diffusion Denoising models demonstrated impressive results across generative Computer Vision tasks, but they still fail to outperform standard autoregressive solutions in the discrete domain, and only match them at best. In this work, we propose a different paradigm by adopting diffusion models to provide suggestions to the autoregressive generation rather than replacing them. By doing so, we combine the bidirectional and refining capabilities of the former with the strong linguistic structure provided by the latter. To showcase its effectiveness, we present Show, Suggest and Tell (SST), which achieves State-of-the-Art results on COCO, among models in a similar setting. In particular, SST achieves 125.1 CIDEr-D on the COCO dataset without Reinforcement Learning, outperforming both autoregressive and diffusion model State-of-the-Art results by 1.5 and 2.5 points. On top of the strong results, we performed extensive experiments to validate the proposal and analyze the impact of the suggestion module. Results demonstrate a positive correlation between suggestion and caption quality, overall indicating a currently underexplored but promising research direction. Code will be available at: https://github.com/jchenghu/show\_suggest\_tell.
title Diffusion Is Your Friend in Show, Suggest and Tell
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2512.10038