Image Generation from Image Captioning -- Invertible Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Menon, Nandakishore S, Kamanchi, Chandramouli, Diddigi, Raghuram Bharadwaj
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914992307044352
author Menon, Nandakishore S
Kamanchi, Chandramouli
Diddigi, Raghuram Bharadwaj
author_facet Menon, Nandakishore S
Kamanchi, Chandramouli
Diddigi, Raghuram Bharadwaj
contents Our work aims to build a model that performs dual tasks of image captioning and image generation while being trained on only one task. The central idea is to train an invertible model that learns a one-to-one mapping between the image and text embeddings. Once the invertible model is efficiently trained on one task, the image captioning, the same model can generate new images for a given text through the inversion process, with no additional training. This paper proposes a simple invertible neural network architecture for this problem and presents our current findings.
format Preprint
id arxiv_https___arxiv_org_abs_2410_20171
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Image Generation from Image Captioning -- Invertible Approach
Menon, Nandakishore S
Kamanchi, Chandramouli
Diddigi, Raghuram Bharadwaj
Computer Vision and Pattern Recognition
Our work aims to build a model that performs dual tasks of image captioning and image generation while being trained on only one task. The central idea is to train an invertible model that learns a one-to-one mapping between the image and text embeddings. Once the invertible model is efficiently trained on one task, the image captioning, the same model can generate new images for a given text through the inversion process, with no additional training. This paper proposes a simple invertible neural network architecture for this problem and presents our current findings.
title Image Generation from Image Captioning -- Invertible Approach
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.20171