Automated Image Captioning with CNNs and Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cahyono, Joshua Adrian, Jusuf, Jeremy Nathan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915063555686400
author Cahyono, Joshua Adrian
Jusuf, Jeremy Nathan
author_facet Cahyono, Joshua Adrian
Jusuf, Jeremy Nathan
contents This project aims to create an automated image captioning system that generates natural language descriptions for input images by integrating techniques from computer vision and natural language processing. We employ various different techniques, ranging from CNN-RNN to the more advanced transformer-based techniques. Training is carried out on image datasets paired with descriptive captions, and model performance will be evaluated using established metrics such as BLEU, METEOR, and CIDEr. The project will also involve experimentation with advanced attention mechanisms, comparisons of different architectural choices, and hyperparameter optimization to refine captioning accuracy and overall system effectiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10511
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Automated Image Captioning with CNNs and Transformers
Cahyono, Joshua Adrian
Jusuf, Jeremy Nathan
Computer Vision and Pattern Recognition
Artificial Intelligence
This project aims to create an automated image captioning system that generates natural language descriptions for input images by integrating techniques from computer vision and natural language processing. We employ various different techniques, ranging from CNN-RNN to the more advanced transformer-based techniques. Training is carried out on image datasets paired with descriptive captions, and model performance will be evaluated using established metrics such as BLEU, METEOR, and CIDEr. The project will also involve experimentation with advanced attention mechanisms, comparisons of different architectural choices, and hyperparameter optimization to refine captioning accuracy and overall system effectiveness.
title Automated Image Captioning with CNNs and Transformers
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2412.10511