LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hashemi, Mohammad Abuzar, Li, Zhanghexuan, Chauhan, Mihir, Shen, Yan, Satbhai, Abhishek, Ali, Mir Basheer, Gao, Mingchen, Srihari, Sargur
Format: Preprint
Veröffentlicht: 2021
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909331535953920
author Hashemi, Mohammad Abuzar
Li, Zhanghexuan
Chauhan, Mihir
Shen, Yan
Satbhai, Abhishek
Ali, Mir Basheer
Gao, Mingchen
Srihari, Sargur
author_facet Hashemi, Mohammad Abuzar
Li, Zhanghexuan
Chauhan, Mihir
Shen, Yan
Satbhai, Abhishek
Ali, Mir Basheer
Gao, Mingchen
Srihari, Sargur
contents Pre-training visual and textual representations from large-scale image-text pairs is becoming a standard approach for many downstream vision-language tasks. The transformer-based models learn inter and intra-modal attention through a list of self-supervised learning tasks. This paper proposes LAViTeR, a novel architecture for visual and textual representation learning. The main module, Visual Textual Alignment (VTA) will be assisted by two auxiliary tasks, GAN-based image synthesis and Image Captioning. We also propose a new evaluation metric measuring the similarity between the learnt visual and textual embedding. The experimental results on two public datasets, CUB and MS-COCO, demonstrate superior visual and textual representation alignment in the joint feature embedding space
format Preprint
id arxiv_https___arxiv_org_abs_2109_04993
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
Hashemi, Mohammad Abuzar
Li, Zhanghexuan
Chauhan, Mihir
Shen, Yan
Satbhai, Abhishek
Ali, Mir Basheer
Gao, Mingchen
Srihari, Sargur
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Pre-training visual and textual representations from large-scale image-text pairs is becoming a standard approach for many downstream vision-language tasks. The transformer-based models learn inter and intra-modal attention through a list of self-supervised learning tasks. This paper proposes LAViTeR, a novel architecture for visual and textual representation learning. The main module, Visual Textual Alignment (VTA) will be assisted by two auxiliary tasks, GAN-based image synthesis and Image Captioning. We also propose a new evaluation metric measuring the similarity between the learnt visual and textual embedding. The experimental results on two public datasets, CUB and MS-COCO, demonstrate superior visual and textual representation alignment in the joint feature embedding space
title LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2109.04993