CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Saito, Kuniaki, Kim, Donghyun, Park, Kwanyong, Hashimoto, Atsushi, Ushiku, Yoshitaka
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911033880215552
author Saito, Kuniaki
Kim, Donghyun
Park, Kwanyong
Hashimoto, Atsushi
Ushiku, Yoshitaka
author_facet Saito, Kuniaki
Kim, Donghyun
Park, Kwanyong
Hashimoto, Atsushi
Ushiku, Yoshitaka
contents An image captioning model flexibly switching its language pattern, e.g., descriptiveness and length, should be useful since it can be applied to diverse applications. However, despite the dramatic improvement in generative vision-language models, fine-grained control over the properties of generated captions is not easy due to two reasons: (i) existing models are not given the properties as a condition during training and (ii) existing models cannot smoothly transition its language pattern from one state to the other. Given this challenge, we propose a new approach, CaptionSmiths, to acquire a single captioning model that can handle diverse language patterns. First, our approach quantifies three properties of each caption, length, descriptiveness, and uniqueness of a word, as continuous scalar values, without human annotation. Given the values, we represent the conditioning via interpolation between two endpoint vectors corresponding to the extreme states, e.g., one for a very short caption and one for a very long caption. Empirical results demonstrate that the resulting model can smoothly change the properties of the output captions and show higher lexical alignment than baselines. For instance, CaptionSmiths reduces the error in controlling caption length by 506\% despite better lexical alignment. Code will be available on https://github.com/omron-sinicx/captionsmiths.
format Preprint
id arxiv_https___arxiv_org_abs_2507_01409
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning
Saito, Kuniaki
Kim, Donghyun
Park, Kwanyong
Hashimoto, Atsushi
Ushiku, Yoshitaka
Computer Vision and Pattern Recognition
An image captioning model flexibly switching its language pattern, e.g., descriptiveness and length, should be useful since it can be applied to diverse applications. However, despite the dramatic improvement in generative vision-language models, fine-grained control over the properties of generated captions is not easy due to two reasons: (i) existing models are not given the properties as a condition during training and (ii) existing models cannot smoothly transition its language pattern from one state to the other. Given this challenge, we propose a new approach, CaptionSmiths, to acquire a single captioning model that can handle diverse language patterns. First, our approach quantifies three properties of each caption, length, descriptiveness, and uniqueness of a word, as continuous scalar values, without human annotation. Given the values, we represent the conditioning via interpolation between two endpoint vectors corresponding to the extreme states, e.g., one for a very short caption and one for a very long caption. Empirical results demonstrate that the resulting model can smoothly change the properties of the output captions and show higher lexical alignment than baselines. For instance, CaptionSmiths reduces the error in controlling caption length by 506\% despite better lexical alignment. Code will be available on https://github.com/omron-sinicx/captionsmiths.
title CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.01409