Guardado en:
Detalles Bibliográficos
Autores principales: Basak, Debolena, Srijith, P. K., Desarkar, Maunendra Sankar
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:https://arxiv.org/abs/2403.06292
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910361797525504
author Basak, Debolena
Srijith, P. K.
Desarkar, Maunendra Sankar
author_facet Basak, Debolena
Srijith, P. K.
Desarkar, Maunendra Sankar
contents In several real-world scenarios like autonomous navigation and mobility, to obtain a better visual understanding of the surroundings, image captioning and object detection play a crucial role. This work introduces a novel multitask learning framework that combines image captioning and object detection into a joint model. We propose TICOD, Transformer-based Image Captioning and Object detection model for jointly training both tasks by combining the losses obtained from image captioning and object detection networks. By leveraging joint training, the model benefits from the complementary information shared between the two tasks, leading to improved performance for image captioning. Our approach utilizes a transformer-based architecture that enables end-to-end network integration for image captioning and object detection and performs both tasks jointly. We evaluate the effectiveness of our approach through comprehensive experiments on the MS-COCO dataset. Our model outperforms the baselines from image captioning literature by achieving a 3.65% improvement in BERTScore.
format Preprint
id arxiv_https___arxiv_org_abs_2403_06292
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Transformer based Multitask Learning for Image Captioning and Object Detection
Basak, Debolena
Srijith, P. K.
Desarkar, Maunendra Sankar
Computer Vision and Pattern Recognition
Computation and Language
In several real-world scenarios like autonomous navigation and mobility, to obtain a better visual understanding of the surroundings, image captioning and object detection play a crucial role. This work introduces a novel multitask learning framework that combines image captioning and object detection into a joint model. We propose TICOD, Transformer-based Image Captioning and Object detection model for jointly training both tasks by combining the losses obtained from image captioning and object detection networks. By leveraging joint training, the model benefits from the complementary information shared between the two tasks, leading to improved performance for image captioning. Our approach utilizes a transformer-based architecture that enables end-to-end network integration for image captioning and object detection and performs both tasks jointly. We evaluate the effectiveness of our approach through comprehensive experiments on the MS-COCO dataset. Our model outperforms the baselines from image captioning literature by achieving a 3.65% improvement in BERTScore.
title Transformer based Multitask Learning for Image Captioning and Object Detection
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2403.06292