FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910869250637824 |
|---|---|
| author | Yun, Jun-Hak Kim, Seung-Bin Lee, Seong-Whan |
| author_facet | Yun, Jun-Hak Kim, Seung-Bin Lee, Seong-Whan |
| contents | Audio super-resolution is challenging owing to its ill-posed nature. Recently, the application of diffusion models in audio super-resolution has shown promising results in alleviating this challenge. However, diffusion-based models have limitations, primarily the necessity for numerous sampling steps, which causes significantly increased latency when synthesizing high-quality audio samples. In this paper, we propose FLowHigh, a novel approach that integrates flow matching, a highly efficient generative model, into audio super-resolution. We also explore probability paths specially tailored for audio super-resolution, which effectively capture high-resolution audio distributions, thereby enhancing reconstruction quality. The proposed method generates high-fidelity, high-resolution audio through a single-step sampling process across various input sampling rates. The experimental results on the VCTK benchmark dataset demonstrate that FLowHigh achieves state-of-the-art performance in audio super-resolution, as evaluated by log-spectral distance and ViSQOL while maintaining computational efficiency with only a single-step sampling process. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_04926 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching Yun, Jun-Hak Kim, Seung-Bin Lee, Seong-Whan Audio and Speech Processing Artificial Intelligence Computation and Language Sound Audio super-resolution is challenging owing to its ill-posed nature. Recently, the application of diffusion models in audio super-resolution has shown promising results in alleviating this challenge. However, diffusion-based models have limitations, primarily the necessity for numerous sampling steps, which causes significantly increased latency when synthesizing high-quality audio samples. In this paper, we propose FLowHigh, a novel approach that integrates flow matching, a highly efficient generative model, into audio super-resolution. We also explore probability paths specially tailored for audio super-resolution, which effectively capture high-resolution audio distributions, thereby enhancing reconstruction quality. The proposed method generates high-fidelity, high-resolution audio through a single-step sampling process across various input sampling rates. The experimental results on the VCTK benchmark dataset demonstrate that FLowHigh achieves state-of-the-art performance in audio super-resolution, as evaluated by log-spectral distance and ViSQOL while maintaining computational efficiency with only a single-step sampling process. |
| title | FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching |
| topic | Audio and Speech Processing Artificial Intelligence Computation and Language Sound |
| url | https://arxiv.org/abs/2501.04926 |