MObyGaze: a film dataset of multimodal objectification densely annotated by experts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tores, Julie, Ancarani, Elisa, Sassatelli, Lucile, Wu, Hui-Yin, Bergman, Clement, Andolfi, Lea, Ecrement, Victor, Sun, Remy, Precioso, Frederic, Devars, Thierry, Guaresi, Magali, Julliard, Virginie, Lecossais, Sarah
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912400019554304
author Tores, Julie
Ancarani, Elisa
Sassatelli, Lucile
Wu, Hui-Yin
Bergman, Clement
Andolfi, Lea
Ecrement, Victor
Sun, Remy
Precioso, Frederic
Devars, Thierry
Guaresi, Magali
Julliard, Virginie
Lecossais, Sarah
author_facet Tores, Julie
Ancarani, Elisa
Sassatelli, Lucile
Wu, Hui-Yin
Bergman, Clement
Andolfi, Lea
Ecrement, Victor
Sun, Remy
Precioso, Frederic
Devars, Thierry
Guaresi, Magali
Julliard, Virginie
Lecossais, Sarah
contents Characterizing and quantifying gender representation disparities in audiovisual storytelling contents is necessary to grasp how stereotypes may perpetuate on screen. In this article, we consider the high-level construct of objectification and introduce a new AI task to the ML community: characterize and quantify complex multimodal (visual, speech, audio) temporal patterns producing objectification in films. Building on film studies and psychology, we define the construct of objectification in a structured thesaurus involving 5 sub-constructs manifesting through 11 concepts spanning 3 modalities. We introduce the Multimodal Objectifying Gaze (MObyGaze) dataset, made of 20 movies annotated densely by experts for objectification levels and concepts over freely delimited segments: it amounts to 6072 segments over 43 hours of video with fine-grained localization and categorization. We formulate different learning tasks, propose and investigate best ways to learn from the diversity of labels among a low number of annotators, and benchmark recent vision, text and audio models, showing the feasibility of the task. We make our code and our dataset available to the community and described in the Croissant format: https://anonymous.4open.science/r/MObyGaze-F600/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22084
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MObyGaze: a film dataset of multimodal objectification densely annotated by experts
Tores, Julie
Ancarani, Elisa
Sassatelli, Lucile
Wu, Hui-Yin
Bergman, Clement
Andolfi, Lea
Ecrement, Victor
Sun, Remy
Precioso, Frederic
Devars, Thierry
Guaresi, Magali
Julliard, Virginie
Lecossais, Sarah
Computer Vision and Pattern Recognition
Characterizing and quantifying gender representation disparities in audiovisual storytelling contents is necessary to grasp how stereotypes may perpetuate on screen. In this article, we consider the high-level construct of objectification and introduce a new AI task to the ML community: characterize and quantify complex multimodal (visual, speech, audio) temporal patterns producing objectification in films. Building on film studies and psychology, we define the construct of objectification in a structured thesaurus involving 5 sub-constructs manifesting through 11 concepts spanning 3 modalities. We introduce the Multimodal Objectifying Gaze (MObyGaze) dataset, made of 20 movies annotated densely by experts for objectification levels and concepts over freely delimited segments: it amounts to 6072 segments over 43 hours of video with fine-grained localization and categorization. We formulate different learning tasks, propose and investigate best ways to learn from the diversity of labels among a low number of annotators, and benchmark recent vision, text and audio models, showing the feasibility of the task. We make our code and our dataset available to the community and described in the Croissant format: https://anonymous.4open.science/r/MObyGaze-F600/.
title MObyGaze: a film dataset of multimodal objectification densely annotated by experts
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.22084