SMILE: Multimodal Dataset for Understanding Laughter in Video with Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hyun, Lee, Sung-Bin, Kim, Han, Seungju, Yu, Youngjae, Oh, Tae-Hyun
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929356033490944
author Hyun, Lee
Sung-Bin, Kim
Han, Seungju
Yu, Youngjae
Oh, Tae-Hyun
author_facet Hyun, Lee
Sung-Bin, Kim
Han, Seungju
Yu, Youngjae
Oh, Tae-Hyun
contents Despite the recent advances of the artificial intelligence, building social intelligence remains a challenge. Among social signals, laughter is one of the distinctive expressions that occurs during social interactions between humans. In this work, we tackle a new challenge for machines to understand the rationale behind laughter in video, Video Laugh Reasoning. We introduce this new task to explain why people laugh in a particular video and a dataset for this task. Our proposed dataset, SMILE, comprises video clips and language descriptions of why people laugh. We propose a baseline by leveraging the reasoning capacity of large language models (LLMs) with textual video representation. Experiments show that our baseline can generate plausible explanations for laughter. We further investigate the scalability of our baseline by probing other video understanding tasks and in-the-wild videos. We release our dataset, code, and model checkpoints on https://github.com/postech-ami/SMILE-Dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2312_09818
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle SMILE: Multimodal Dataset for Understanding Laughter in Video with Language Models
Hyun, Lee
Sung-Bin, Kim
Han, Seungju
Yu, Youngjae
Oh, Tae-Hyun
Computation and Language
Artificial Intelligence
Despite the recent advances of the artificial intelligence, building social intelligence remains a challenge. Among social signals, laughter is one of the distinctive expressions that occurs during social interactions between humans. In this work, we tackle a new challenge for machines to understand the rationale behind laughter in video, Video Laugh Reasoning. We introduce this new task to explain why people laugh in a particular video and a dataset for this task. Our proposed dataset, SMILE, comprises video clips and language descriptions of why people laugh. We propose a baseline by leveraging the reasoning capacity of large language models (LLMs) with textual video representation. Experiments show that our baseline can generate plausible explanations for laughter. We further investigate the scalability of our baseline by probing other video understanding tasks and in-the-wild videos. We release our dataset, code, and model checkpoints on https://github.com/postech-ami/SMILE-Dataset.
title SMILE: Multimodal Dataset for Understanding Laughter in Video with Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2312.09818