LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2021
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866909331535953920 |
|---|---|
| author | Hashemi, Mohammad Abuzar Li, Zhanghexuan Chauhan, Mihir Shen, Yan Satbhai, Abhishek Ali, Mir Basheer Gao, Mingchen Srihari, Sargur |
| author_facet | Hashemi, Mohammad Abuzar Li, Zhanghexuan Chauhan, Mihir Shen, Yan Satbhai, Abhishek Ali, Mir Basheer Gao, Mingchen Srihari, Sargur |
| contents | Pre-training visual and textual representations from large-scale image-text pairs is becoming a standard approach for many downstream vision-language tasks. The transformer-based models learn inter and intra-modal attention through a list of self-supervised learning tasks. This paper proposes LAViTeR, a novel architecture for visual and textual representation learning. The main module, Visual Textual Alignment (VTA) will be assisted by two auxiliary tasks, GAN-based image synthesis and Image Captioning. We also propose a new evaluation metric measuring the similarity between the learnt visual and textual embedding. The experimental results on two public datasets, CUB and MS-COCO, demonstrate superior visual and textual representation alignment in the joint feature embedding space |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2109_04993 |
| institution | arXiv |
| publishDate | 2021 |
| record_format | arxiv |
| spellingShingle | LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation Hashemi, Mohammad Abuzar Li, Zhanghexuan Chauhan, Mihir Shen, Yan Satbhai, Abhishek Ali, Mir Basheer Gao, Mingchen Srihari, Sargur Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Pre-training visual and textual representations from large-scale image-text pairs is becoming a standard approach for many downstream vision-language tasks. The transformer-based models learn inter and intra-modal attention through a list of self-supervised learning tasks. This paper proposes LAViTeR, a novel architecture for visual and textual representation learning. The main module, Visual Textual Alignment (VTA) will be assisted by two auxiliary tasks, GAN-based image synthesis and Image Captioning. We also propose a new evaluation metric measuring the similarity between the learnt visual and textual embedding. The experimental results on two public datasets, CUB and MS-COCO, demonstrate superior visual and textual representation alignment in the joint feature embedding space |
| title | LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2109.04993 |