Unboxing Engagement in YouTube Influencer Videos: An Attention-Based Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rajaram, Prashant, Manchanda, Puneet
Format: Preprint
Published: 2020
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912368828612608
author Rajaram, Prashant
Manchanda, Puneet
author_facet Rajaram, Prashant
Manchanda, Puneet
contents Influencer marketing has become a widely used strategy for reaching customers. Despite growing interest among influencers and brand partners in predicting engagement with influencer videos, there has been little research on the relative importance of different video data modalities in predicting engagement. We analyze unstructured data from long-form YouTube influencer videos - spanning text, audio, and video images - using an interpretable deep learning framework that leverages model attention to video elements. This framework enables strong out-of-sample prediction, followed by ex-post interpretation using a novel approach that prunes spurious associations. Our prediction-based results reveal that "what is said" through words (text) is more important than "how it is said" through imagery (video images) or acoustics (audio) in predicting video engagement. Interpretation-based findings show that during the critical onset period of a video (first 30 seconds), auditory stimuli (e.g., brand mentions and music) are associated with sentiment expressed in verbal engagement (comments), while visual stimuli (e.g., video images of humans and packaged goods) are linked with sentiment expressed through non-verbal engagement (the thumbs-up/down ratio). We validate our approach through multiple methods, connect our findings to relevant theory, and discuss implications for influencers, brands and agencies.
format Preprint
id arxiv_https___arxiv_org_abs_2012_12311
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle Unboxing Engagement in YouTube Influencer Videos: An Attention-Based Approach
Rajaram, Prashant
Manchanda, Puneet
Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
Influencer marketing has become a widely used strategy for reaching customers. Despite growing interest among influencers and brand partners in predicting engagement with influencer videos, there has been little research on the relative importance of different video data modalities in predicting engagement. We analyze unstructured data from long-form YouTube influencer videos - spanning text, audio, and video images - using an interpretable deep learning framework that leverages model attention to video elements. This framework enables strong out-of-sample prediction, followed by ex-post interpretation using a novel approach that prunes spurious associations. Our prediction-based results reveal that "what is said" through words (text) is more important than "how it is said" through imagery (video images) or acoustics (audio) in predicting video engagement. Interpretation-based findings show that during the critical onset period of a video (first 30 seconds), auditory stimuli (e.g., brand mentions and music) are associated with sentiment expressed in verbal engagement (comments), while visual stimuli (e.g., video images of humans and packaged goods) are linked with sentiment expressed through non-verbal engagement (the thumbs-up/down ratio). We validate our approach through multiple methods, connect our findings to relevant theory, and discuss implications for influencers, brands and agencies.
title Unboxing Engagement in YouTube Influencer Videos: An Attention-Based Approach
topic Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2012.12311