MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ning, Zheng, Zhang, Zheng, Ban, Jerrick, Jiang, Kaiwen, Gan, Ruohong, Tian, Yapeng, Li, Toby Jia-Jun
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910419343376384
author Ning, Zheng
Zhang, Zheng
Ban, Jerrick
Jiang, Kaiwen
Gan, Ruohong
Tian, Yapeng
Li, Toby Jia-Jun
author_facet Ning, Zheng
Zhang, Zheng
Ban, Jerrick
Jiang, Kaiwen
Gan, Ruohong
Tian, Yapeng
Li, Toby Jia-Jun
contents Spatial audio offers more immersive video consumption experiences to viewers; however, creating and editing spatial audio often expensive and requires specialized equipment and skills, posing a high barrier for amateur video creators. We present MIMOSA, a human-AI co-creation tool that enables amateur users to computationally generate and manipulate spatial audio effects. For a video with only monaural or stereo audio, MIMOSA automatically grounds each sound source to the corresponding sounding object in the visual scene and enables users to further validate and fix the errors in the locations of sounding objects. Users can also augment the spatial audio effect by flexibly manipulating the sounding source positions and creatively customizing the audio effect. The design of MIMOSA exemplifies a human-AI collaboration approach that, instead of utilizing state-of art end-to-end "black-box" ML models, uses a multistep pipeline that aligns its interpretable intermediate results with the user's workflow. A lab user study with 15 participants demonstrates MIMOSA's usability, usefulness, expressiveness, and capability in creating immersive spatial audio effects in collaboration with users.
format Preprint
id arxiv_https___arxiv_org_abs_2404_15107
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos
Ning, Zheng
Zhang, Zheng
Ban, Jerrick
Jiang, Kaiwen
Gan, Ruohong
Tian, Yapeng
Li, Toby Jia-Jun
Human-Computer Interaction
Multimedia
Spatial audio offers more immersive video consumption experiences to viewers; however, creating and editing spatial audio often expensive and requires specialized equipment and skills, posing a high barrier for amateur video creators. We present MIMOSA, a human-AI co-creation tool that enables amateur users to computationally generate and manipulate spatial audio effects. For a video with only monaural or stereo audio, MIMOSA automatically grounds each sound source to the corresponding sounding object in the visual scene and enables users to further validate and fix the errors in the locations of sounding objects. Users can also augment the spatial audio effect by flexibly manipulating the sounding source positions and creatively customizing the audio effect. The design of MIMOSA exemplifies a human-AI collaboration approach that, instead of utilizing state-of art end-to-end "black-box" ML models, uses a multistep pipeline that aligns its interpretable intermediate results with the user's workflow. A lab user study with 15 participants demonstrates MIMOSA's usability, usefulness, expressiveness, and capability in creating immersive spatial audio effects in collaboration with users.
title MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos
topic Human-Computer Interaction
Multimedia
url https://arxiv.org/abs/2404.15107