Saved in:
Bibliographic Details
Main Authors: Haydarov, Kilichbek, Shen, Xiaoqian, Madasu, Avinash, Salem, Mahmoud, Li, Li-Jia, Elsayed, Gamaleldin, Elhoseiny, Mohamed
Format: Preprint
Published: 2023
Subjects:
Online Access:https://arxiv.org/abs/2308.16349
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913630288609280
author Haydarov, Kilichbek
Shen, Xiaoqian
Madasu, Avinash
Salem, Mahmoud
Li, Li-Jia
Elsayed, Gamaleldin
Elhoseiny, Mohamed
author_facet Haydarov, Kilichbek
Shen, Xiaoqian
Madasu, Avinash
Salem, Mahmoud
Li, Li-Jia
Elsayed, Gamaleldin
Elhoseiny, Mohamed
contents We introduce Affective Visual Dialog, an emotion explanation and reasoning task as a testbed for research on understanding the formation of emotions in visually grounded conversations. The task involves three skills: (1) Dialog-based Question Answering (2) Dialog-based Emotion Prediction and (3) Affective emotion explanation generation based on the dialog. Our key contribution is the collection of a large-scale dataset, dubbed AffectVisDial, consisting of 50K 10-turn visually grounded dialogs as well as concluding emotion attributions and dialog-informed textual emotion explanations, resulting in a total of 27,180 working hours. We explain our design decisions in collecting the dataset and introduce the questioner and answerer tasks that are associated with the participants in the conversation. We train and demonstrate solid Affective Visual Dialog baselines adapted from state-of-the-art models. Remarkably, the responses generated by our models show promising emotional reasoning abilities in response to visually grounded conversations. Our project page is available at https://affective-visual-dialog.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2308_16349
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations
Haydarov, Kilichbek
Shen, Xiaoqian
Madasu, Avinash
Salem, Mahmoud
Li, Li-Jia
Elsayed, Gamaleldin
Elhoseiny, Mohamed
Computation and Language
We introduce Affective Visual Dialog, an emotion explanation and reasoning task as a testbed for research on understanding the formation of emotions in visually grounded conversations. The task involves three skills: (1) Dialog-based Question Answering (2) Dialog-based Emotion Prediction and (3) Affective emotion explanation generation based on the dialog. Our key contribution is the collection of a large-scale dataset, dubbed AffectVisDial, consisting of 50K 10-turn visually grounded dialogs as well as concluding emotion attributions and dialog-informed textual emotion explanations, resulting in a total of 27,180 working hours. We explain our design decisions in collecting the dataset and introduce the questioner and answerer tasks that are associated with the participants in the conversation. We train and demonstrate solid Affective Visual Dialog baselines adapted from state-of-the-art models. Remarkably, the responses generated by our models show promising emotional reasoning abilities in response to visually grounded conversations. Our project page is available at https://affective-visual-dialog.github.io.
title Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations
topic Computation and Language
url https://arxiv.org/abs/2308.16349