Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Che, Chang, Lin, Qunwei, Zhao, Xinyu, Huang, Jiaxin, Yu, Liqiang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911755920211968
author Che, Chang
Lin, Qunwei
Zhao, Xinyu
Huang, Jiaxin
Yu, Liqiang
author_facet Che, Chang
Lin, Qunwei
Zhao, Xinyu
Huang, Jiaxin
Yu, Liqiang
contents The process of transforming input images into corresponding textual explanations stands as a crucial and complex endeavor within the domains of computer vision and natural language processing. In this paper, we propose an innovative ensemble approach that harnesses the capabilities of Contrastive Language-Image Pretraining models.
format Preprint
id arxiv_https___arxiv_org_abs_2401_06167
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation
Che, Chang
Lin, Qunwei
Zhao, Xinyu
Huang, Jiaxin
Yu, Liqiang
Computer Vision and Pattern Recognition
Artificial Intelligence
The process of transforming input images into corresponding textual explanations stands as a crucial and complex endeavor within the domains of computer vision and natural language processing. In this paper, we propose an innovative ensemble approach that harnesses the capabilities of Contrastive Language-Image Pretraining models.
title Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2401.06167