Saved in:
Bibliographic Details
Main Authors: Lee, Jaejun, Oh, Yoori, Lee, Kyogu
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.01879
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • In this paper, we introduce a novel framework for generating multi-speaker speech without relying on any audible inputs. Our approach leverages silent electromyography (EMG) signals to capture linguistic content, while facial images are used to match with the vocal identity of the target speaker. Notably, we present a pitch-disentangled content embedding that enhances the extraction of linguistic content from EMG signals. Extensive analysis demonstrates that our method can generate multi-speaker speech without any audible inputs and confirms the effectiveness of the proposed pitch-disentanglement approach.