MMAudio-LABEL: Audio Event Labeling via Audio Generation for Silent Video

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tateishi, Kazuya, Takahashi, Akira, Hiroe, Atsuo, Takeda, Hirofumi, Takahashi, Shusuke, Mitsufuji, Yuki
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917452930088960
author Tateishi, Kazuya
Takahashi, Akira
Hiroe, Atsuo
Takeda, Hirofumi
Takahashi, Shusuke
Mitsufuji, Yuki
author_facet Tateishi, Kazuya
Takahashi, Akira
Hiroe, Atsuo
Takeda, Hirofumi
Takahashi, Shusuke
Mitsufuji, Yuki
contents Recent advances in multimodal generation have enabled high-quality audio generation from silent videos. Practical applications, such as sound production, demand not only the generated audio but also explicit sound event labels detailing the type and timing of sounds. One straightforward approach involves applying a standard sound event detection to the generated audio. However, this post-hoc pipeline is inherently limited, as it is prone to error accumulation. To address this limitation, we propose MMAudio-LABEL (LAtent-Based Event Labeling), an event-aware audio generation framework built on a foundational audio generation model as its backbone that jointly generates audio and frame-aligned sound event predictions from silent videos. We evaluate our method on the Greatest Hits dataset for onset detection and 17-class material classification. Our approach improves onset-detection accuracy from 46.7% to 75.0% and material-classification accuracy from 40.6% to 61.0% over baselines. These results suggest that jointly learning audio generation and event prediction enables a more interpretable and practical video-to-audio synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2605_00495
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MMAudio-LABEL: Audio Event Labeling via Audio Generation for Silent Video
Tateishi, Kazuya
Takahashi, Akira
Hiroe, Atsuo
Takeda, Hirofumi
Takahashi, Shusuke
Mitsufuji, Yuki
Sound
Computer Vision and Pattern Recognition
Recent advances in multimodal generation have enabled high-quality audio generation from silent videos. Practical applications, such as sound production, demand not only the generated audio but also explicit sound event labels detailing the type and timing of sounds. One straightforward approach involves applying a standard sound event detection to the generated audio. However, this post-hoc pipeline is inherently limited, as it is prone to error accumulation. To address this limitation, we propose MMAudio-LABEL (LAtent-Based Event Labeling), an event-aware audio generation framework built on a foundational audio generation model as its backbone that jointly generates audio and frame-aligned sound event predictions from silent videos. We evaluate our method on the Greatest Hits dataset for onset detection and 17-class material classification. Our approach improves onset-detection accuracy from 46.7% to 75.0% and material-classification accuracy from 40.6% to 61.0% over baselines. These results suggest that jointly learning audio generation and event prediction enables a more interpretable and practical video-to-audio synthesis.
title MMAudio-LABEL: Audio Event Labeling via Audio Generation for Silent Video
topic Sound
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.00495