UniFlow-Audio: Unified Flow Matching for Audio Generation from Omni-Modalities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Xuenan, Mei, Jiahao, Zheng, Zihao, Tao, Ye, Xie, Zeyu, Zhang, Yaoyun, Liu, Haohe, Wu, Yuning, Yan, Ming, Wu, Wen, Zhang, Chao, Wu, Mengyue
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916976140484608
author Xu, Xuenan
Mei, Jiahao
Zheng, Zihao
Tao, Ye
Xie, Zeyu
Zhang, Yaoyun
Liu, Haohe
Wu, Yuning
Yan, Ming
Wu, Wen
Zhang, Chao
Wu, Mengyue
author_facet Xu, Xuenan
Mei, Jiahao
Zheng, Zihao
Tao, Ye
Xie, Zeyu
Zhang, Yaoyun
Liu, Haohe
Wu, Yuning
Yan, Ming
Wu, Wen
Zhang, Chao
Wu, Mengyue
contents Audio generation, including speech, music and sound effects, has advanced rapidly in recent years. These tasks can be divided into two categories: time-aligned (TA) tasks, where each input unit corresponds to a specific segment of the output audio (e.g., phonemes aligned with frames in speech synthesis); and non-time-aligned (NTA) tasks, where such alignment is not available. Since modeling paradigms for the two types are typically different, research on different audio generation tasks has traditionally followed separate trajectories. However, audio is not inherently divided into such categories, making a unified model a natural and necessary goal for general audio generation. Previous unified audio generation works have adopted autoregressive architectures, while unified non-autoregressive approaches remain largely unexplored. In this work, we propose UniFlow-Audio, a universal audio generation framework based on flow matching. We propose a dual-fusion mechanism that temporally aligns audio latents with TA features and integrates NTA features via cross-attention in each model block. Task-balanced data sampling is employed to maintain strong performance across both TA and NTA tasks. UniFlow-Audio supports omni-modalities, including text, audio, and video. By leveraging the advantage of multi-task learning and the generative modeling capabilities of flow matching, UniFlow-Audio achieves strong results across 7 tasks using fewer than 8K hours of public training data and under 1B trainable parameters. Even the small variant with only ~200M trainable parameters shows competitive performance, highlighting UniFlow-Audio as a potential non-auto-regressive foundation model for audio generation. Code and models will be available at https://wsntxxn.github.io/uniflow_audio.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24391
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UniFlow-Audio: Unified Flow Matching for Audio Generation from Omni-Modalities
Xu, Xuenan
Mei, Jiahao
Zheng, Zihao
Tao, Ye
Xie, Zeyu
Zhang, Yaoyun
Liu, Haohe
Wu, Yuning
Yan, Ming
Wu, Wen
Zhang, Chao
Wu, Mengyue
Sound
Audio generation, including speech, music and sound effects, has advanced rapidly in recent years. These tasks can be divided into two categories: time-aligned (TA) tasks, where each input unit corresponds to a specific segment of the output audio (e.g., phonemes aligned with frames in speech synthesis); and non-time-aligned (NTA) tasks, where such alignment is not available. Since modeling paradigms for the two types are typically different, research on different audio generation tasks has traditionally followed separate trajectories. However, audio is not inherently divided into such categories, making a unified model a natural and necessary goal for general audio generation. Previous unified audio generation works have adopted autoregressive architectures, while unified non-autoregressive approaches remain largely unexplored. In this work, we propose UniFlow-Audio, a universal audio generation framework based on flow matching. We propose a dual-fusion mechanism that temporally aligns audio latents with TA features and integrates NTA features via cross-attention in each model block. Task-balanced data sampling is employed to maintain strong performance across both TA and NTA tasks. UniFlow-Audio supports omni-modalities, including text, audio, and video. By leveraging the advantage of multi-task learning and the generative modeling capabilities of flow matching, UniFlow-Audio achieves strong results across 7 tasks using fewer than 8K hours of public training data and under 1B trainable parameters. Even the small variant with only ~200M trainable parameters shows competitive performance, highlighting UniFlow-Audio as a potential non-auto-regressive foundation model for audio generation. Code and models will be available at https://wsntxxn.github.io/uniflow_audio.
title UniFlow-Audio: Unified Flow Matching for Audio Generation from Omni-Modalities
topic Sound
url https://arxiv.org/abs/2509.24391