Holistic Visual-Textual Sentiment Analysis with Prior Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Junyu, An, Jie, Lyu, Hanjia, Kanan, Christopher, Luo, Jiebo
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913382757564416
author Chen, Junyu
An, Jie
Lyu, Hanjia
Kanan, Christopher
Luo, Jiebo
author_facet Chen, Junyu
An, Jie
Lyu, Hanjia
Kanan, Christopher
Luo, Jiebo
contents Visual-textual sentiment analysis aims to predict sentiment with the input of a pair of image and text, which poses a challenge in learning effective features for diverse input images. To address this, we propose a holistic method that achieves robust visual-textual sentiment analysis by exploiting a rich set of powerful pre-trained visual and textual prior models. The proposed method consists of four parts: (1) a visual-textual branch to learn features directly from data for sentiment analysis, (2) a visual expert branch with a set of pre-trained "expert" encoders to extract selected semantic visual features, (3) a CLIP branch to implicitly model visual-textual correspondence, and (4) a multimodal feature fusion network based on BERT to fuse multimodal features and make sentiment predictions. Extensive experiments on three datasets show that our method produces better visual-textual sentiment analysis performance than existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2211_12981
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Holistic Visual-Textual Sentiment Analysis with Prior Models
Chen, Junyu
An, Jie
Lyu, Hanjia
Kanan, Christopher
Luo, Jiebo
Computer Vision and Pattern Recognition
Multimedia
Visual-textual sentiment analysis aims to predict sentiment with the input of a pair of image and text, which poses a challenge in learning effective features for diverse input images. To address this, we propose a holistic method that achieves robust visual-textual sentiment analysis by exploiting a rich set of powerful pre-trained visual and textual prior models. The proposed method consists of four parts: (1) a visual-textual branch to learn features directly from data for sentiment analysis, (2) a visual expert branch with a set of pre-trained "expert" encoders to extract selected semantic visual features, (3) a CLIP branch to implicitly model visual-textual correspondence, and (4) a multimodal feature fusion network based on BERT to fuse multimodal features and make sentiment predictions. Extensive experiments on three datasets show that our method produces better visual-textual sentiment analysis performance than existing methods.
title Holistic Visual-Textual Sentiment Analysis with Prior Models
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2211.12981