Bone-conduction Guided Multimodal Speech Enhancement with Conditional Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khanagha, Sina, Lay, Bunlong, Gerkmann, Timo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909994177265664
author Khanagha, Sina
Lay, Bunlong
Gerkmann, Timo
author_facet Khanagha, Sina
Lay, Bunlong
Gerkmann, Timo
contents Single-channel speech enhancement models face significant performance degradation in extremely noisy environments. While prior work has shown that complementary bone-conducted speech can guide enhancement, effective integration of this noise-immune modality remains a challenge. This paper introduces a novel multimodal speech enhancement framework that integrates bone-conduction sensors with air-conducted microphones using a conditional diffusion model. Our proposed model significantly outperforms previously established multimodal techniques and a powerful diffusion-based single-modal baseline across a wide range of acoustic conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2601_12354
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Bone-conduction Guided Multimodal Speech Enhancement with Conditional Diffusion Models
Khanagha, Sina
Lay, Bunlong
Gerkmann, Timo
Audio and Speech Processing
Machine Learning
Sound
Single-channel speech enhancement models face significant performance degradation in extremely noisy environments. While prior work has shown that complementary bone-conducted speech can guide enhancement, effective integration of this noise-immune modality remains a challenge. This paper introduces a novel multimodal speech enhancement framework that integrates bone-conduction sensors with air-conducted microphones using a conditional diffusion model. Our proposed model significantly outperforms previously established multimodal techniques and a powerful diffusion-based single-modal baseline across a wide range of acoustic conditions.
title Bone-conduction Guided Multimodal Speech Enhancement with Conditional Diffusion Models
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2601.12354