Large Language Models Meet Text-Centric Multimodal Sentiment Analysis: A Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Hao, Zhao, Yanyan, Wu, Yang, Wang, Shilong, Zheng, Tian, Zhang, Hongbo, Ma, Zongyang, Che, Wanxiang, Qin, Bing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929461535965184
author Yang, Hao
Zhao, Yanyan
Wu, Yang
Wang, Shilong
Zheng, Tian
Zhang, Hongbo
Ma, Zongyang
Che, Wanxiang
Qin, Bing
author_facet Yang, Hao
Zhao, Yanyan
Wu, Yang
Wang, Shilong
Zheng, Tian
Zhang, Hongbo
Ma, Zongyang
Che, Wanxiang
Qin, Bing
contents Compared to traditional sentiment analysis, which only considers text, multimodal sentiment analysis needs to consider emotional signals from multimodal sources simultaneously and is therefore more consistent with the way how humans process sentiment in real-world scenarios. It involves processing emotional information from various sources such as natural language, images, videos, audio, physiological signals, etc. However, although other modalities also contain diverse emotional cues, natural language usually contains richer contextual information and therefore always occupies a crucial position in multimodal sentiment analysis. The emergence of ChatGPT has opened up immense potential for applying large language models (LLMs) to text-centric multimodal tasks. However, it is still unclear how existing LLMs can adapt better to text-centric multimodal sentiment analysis tasks. This survey aims to (1) present a comprehensive review of recent research in text-centric multimodal sentiment analysis tasks, (2) examine the potential of LLMs for text-centric multimodal sentiment analysis, outlining their approaches, advantages, and limitations, (3) summarize the application scenarios of LLM-based multimodal sentiment analysis technology, and (4) explore the challenges and potential research directions for multimodal sentiment analysis in the future.
format Preprint
id arxiv_https___arxiv_org_abs_2406_08068
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Large Language Models Meet Text-Centric Multimodal Sentiment Analysis: A Survey
Yang, Hao
Zhao, Yanyan
Wu, Yang
Wang, Shilong
Zheng, Tian
Zhang, Hongbo
Ma, Zongyang
Che, Wanxiang
Qin, Bing
Computation and Language
Compared to traditional sentiment analysis, which only considers text, multimodal sentiment analysis needs to consider emotional signals from multimodal sources simultaneously and is therefore more consistent with the way how humans process sentiment in real-world scenarios. It involves processing emotional information from various sources such as natural language, images, videos, audio, physiological signals, etc. However, although other modalities also contain diverse emotional cues, natural language usually contains richer contextual information and therefore always occupies a crucial position in multimodal sentiment analysis. The emergence of ChatGPT has opened up immense potential for applying large language models (LLMs) to text-centric multimodal tasks. However, it is still unclear how existing LLMs can adapt better to text-centric multimodal sentiment analysis tasks. This survey aims to (1) present a comprehensive review of recent research in text-centric multimodal sentiment analysis tasks, (2) examine the potential of LLMs for text-centric multimodal sentiment analysis, outlining their approaches, advantages, and limitations, (3) summarize the application scenarios of LLM-based multimodal sentiment analysis technology, and (4) explore the challenges and potential research directions for multimodal sentiment analysis in the future.
title Large Language Models Meet Text-Centric Multimodal Sentiment Analysis: A Survey
topic Computation and Language
url https://arxiv.org/abs/2406.08068