Alternating Approach-Putt Models for Multi-Stage Speech Enhancement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jeong, Iksoon, Kim, Kyung-Joong, Ahn, Kang-Hun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911105519976448
author Jeong, Iksoon
Kim, Kyung-Joong
Ahn, Kang-Hun
author_facet Jeong, Iksoon
Kim, Kyung-Joong
Ahn, Kang-Hun
contents Speech enhancement using artificial neural networks aims to remove noise from noisy speech signals while preserving the speech content. However, speech enhancement networks often introduce distortions to the speech signal, referred to as artifacts, which can degrade audio quality. In this work, we propose a post-processing neural network designed to mitigate artifacts introduced by speech enhancement models. Inspired by the analogy of making a `Putt' after an `Approach' in golf, we name our model PuttNet. We demonstrate that alternating between a speech enhancement model and the proposed Putt model leads to improved speech quality, as measured by perceptual quality scores (PESQ), objective intelligibility (STOI), and background noise intrusiveness (CBAK) scores. Furthermore, we illustrate with graphical analysis why this alternating Approach outperforms repeated application of either model alone.
format Preprint
id arxiv_https___arxiv_org_abs_2508_10436
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Alternating Approach-Putt Models for Multi-Stage Speech Enhancement
Jeong, Iksoon
Kim, Kyung-Joong
Ahn, Kang-Hun
Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
Speech enhancement using artificial neural networks aims to remove noise from noisy speech signals while preserving the speech content. However, speech enhancement networks often introduce distortions to the speech signal, referred to as artifacts, which can degrade audio quality. In this work, we propose a post-processing neural network designed to mitigate artifacts introduced by speech enhancement models. Inspired by the analogy of making a `Putt' after an `Approach' in golf, we name our model PuttNet. We demonstrate that alternating between a speech enhancement model and the proposed Putt model leads to improved speech quality, as measured by perceptual quality scores (PESQ), objective intelligibility (STOI), and background noise intrusiveness (CBAK) scores. Furthermore, we illustrate with graphical analysis why this alternating Approach outperforms repeated application of either model alone.
title Alternating Approach-Putt Models for Multi-Stage Speech Enhancement
topic Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2508.10436