Lightweight and perceptually-guided voice conversion for electro-laryngeal speech
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Mayrhofer, Benedikt, Pernkopf, Franz, Aichinger, Philipp, Hagmüller, Martin |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
A multimodal Bayesian Network for symptom-level depression and anxiety prediction from voice and speech data
par: Norbury, Agnes, et autres
Publié: (2025)
par: Norbury, Agnes, et autres
Publié: (2025)
CognoSpeak: an automatic, remote assessment of early cognitive decline in real-world conversational speech
par: Pahar, Madhurananda, et autres
Publié: (2025)
par: Pahar, Madhurananda, et autres
Publié: (2025)
Online speaker diarization of meetings guided by speech separation
par: Gruttadauria, Elio, et autres
Publié: (2024)
par: Gruttadauria, Elio, et autres
Publié: (2024)
voice2mode: Phonation Mode Classification in Singing using Self-Supervised Speech Models
par: Justus, Aju Ani, et autres
Publié: (2026)
par: Justus, Aju Ani, et autres
Publié: (2026)
Acoustic and perceptual differences between standard and accented speech and their voice clones
par: Yang, Tianle, et autres
Publié: (2026)
par: Yang, Tianle, et autres
Publié: (2026)
IsoNet: Spatially-aware audio-visual target speech extraction in complex acoustic environments
par: Padhya, Dinanath, et autres
Publié: (2026)
par: Padhya, Dinanath, et autres
Publié: (2026)
Throat and acoustic paired speech dataset for deep learning-based speech enhancement
par: Kim, Yunsik, et autres
Publié: (2025)
par: Kim, Yunsik, et autres
Publié: (2025)
Resource-constrained stereo singing voice cancellation
par: Borrelli, Clara, et autres
Publié: (2024)
par: Borrelli, Clara, et autres
Publié: (2024)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
par: Maiti, Soumi, et autres
Publié: (2023)
par: Maiti, Soumi, et autres
Publié: (2023)
Selfsupervised learning for pathological speech detection
par: Sheikh, Shakeel Ahmad
Publié: (2024)
par: Sheikh, Shakeel Ahmad
Publié: (2024)
Towards the Synthesis of Non-speech Vocalizations
par: Hoq, Enjamamul, et autres
Publié: (2024)
par: Hoq, Enjamamul, et autres
Publié: (2024)
Voxceleb-ESP: preliminary experiments detecting Spanish celebrities from their voices
par: Labrador, Beltrán, et autres
Publié: (2023)
par: Labrador, Beltrán, et autres
Publié: (2023)
SpectroFusion-ViT: A Lightweight Transformer for Speech Emotion Recognition Using Harmonic Mel-Chroma Fusion
par: Ahmed, Faria, et autres
Publié: (2026)
par: Ahmed, Faria, et autres
Publié: (2026)
EEG-to-Voice Decoding of Spoken and Imagined speech Using Non-Invasive EEG
par: Park, Hanbeot, et autres
Publié: (2025)
par: Park, Hanbeot, et autres
Publié: (2025)
Beyond saliency: enhancing explanation of speech emotion recognition with expert-referenced acoustic cues
par: Nasr, Seham, et autres
Publié: (2025)
par: Nasr, Seham, et autres
Publié: (2025)
Introduction to speech recognition
par: Dauphin, Gabriel
Publié: (2024)
par: Dauphin, Gabriel
Publié: (2024)
Single-channel speech enhancement using learnable loss mixup
par: Chang, Oscar, et autres
Publié: (2023)
par: Chang, Oscar, et autres
Publié: (2023)
Zipformer: A faster and better encoder for automatic speech recognition
par: Yao, Zengwei, et autres
Publié: (2023)
par: Yao, Zengwei, et autres
Publié: (2023)
CR-CTC: Consistency regularization on CTC for improved speech recognition
par: Yao, Zengwei, et autres
Publié: (2024)
par: Yao, Zengwei, et autres
Publié: (2024)
Robustifying automatic speech recognition by extracting slowly varying features
par: Pizarro, Matías, et autres
Publié: (2021)
par: Pizarro, Matías, et autres
Publié: (2021)
Generalizable speech deepfake detection via meta-learned LoRA
par: Laakkonen, Janne, et autres
Publié: (2025)
par: Laakkonen, Janne, et autres
Publié: (2025)
Late fusion ensembles for speech recognition on diverse input audio representations
par: Jezidžić, Marin, et autres
Publié: (2024)
par: Jezidžić, Marin, et autres
Publié: (2024)
Boosting keyword spotting through on-device learnable user speech characteristics
par: Cioflan, Cristian, et autres
Publié: (2024)
par: Cioflan, Cristian, et autres
Publié: (2024)
Context-aware child-directed speech detection from long-form recordings
par: Charlot, Théo, et autres
Publié: (2026)
par: Charlot, Théo, et autres
Publié: (2026)
Dementia classification from spontaneous speech using wrapper-based feature selection
par: Niemelä, Marko, et autres
Publié: (2025)
par: Niemelä, Marko, et autres
Publié: (2025)
Acoustic characterization of speech rhythm: going beyond metrics with recurrent neural networks
par: Deloche, François, et autres
Publié: (2024)
par: Deloche, François, et autres
Publié: (2024)
FlowDec: A flow-based full-band general audio codec with high perceptual quality
par: Welker, Simon, et autres
Publié: (2025)
par: Welker, Simon, et autres
Publié: (2025)
Fusion approaches for emotion recognition from speech using acoustic and text-based features
par: Pepino, Leonardo, et autres
Publié: (2024)
par: Pepino, Leonardo, et autres
Publié: (2024)
An Attention Long Short-Term Memory based system for automatic classification of speech intelligibility
par: Fernández-Díaz, Miguel, et autres
Publié: (2024)
par: Fernández-Díaz, Miguel, et autres
Publié: (2024)
SeMaScore : a new evaluation metric for automatic speech recognition tasks
par: Sasindran, Zitha, et autres
Publié: (2024)
par: Sasindran, Zitha, et autres
Publié: (2024)
Acoustic-to-articulatory inversion for dysarthric speech: Are pre-trained self-supervised representations favorable?
par: Maharana, Sarthak Kumar, et autres
Publié: (2023)
par: Maharana, Sarthak Kumar, et autres
Publié: (2023)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
par: Fujita, Kenichi, et autres
Publié: (2024)
par: Fujita, Kenichi, et autres
Publié: (2024)
Analyzing the Importance of Blank for CTC-Based Knowledge Distillation
par: Hilmes, Benedikt, et autres
Publié: (2025)
par: Hilmes, Benedikt, et autres
Publié: (2025)
Screening method for early dementia using sound objects as voice biomarkers
par: Pluta, Adam, et autres
Publié: (2024)
par: Pluta, Adam, et autres
Publié: (2024)
Towards objective and interpretable speech disorder assessment: a comparative analysis of CNN and transformer-based models
par: Maisonneuve, Malo, et autres
Publié: (2024)
par: Maisonneuve, Malo, et autres
Publié: (2024)
Objective and subjective evaluation of speech enhancement methods in the UDASE task of the 7th CHiME challenge
par: Leglaive, Simon, et autres
Publié: (2024)
par: Leglaive, Simon, et autres
Publié: (2024)
A multimodal dynamical variational autoencoder for audiovisual speech representation learning
par: Sadok, Samir, et autres
Publié: (2023)
par: Sadok, Samir, et autres
Publié: (2023)
A vector quantized masked autoencoder for audiovisual speech emotion recognition
par: Sadok, Samir, et autres
Publié: (2023)
par: Sadok, Samir, et autres
Publié: (2023)
A comparative study of generative models for child voice conversion
par: Sudro, Protima Nomo, et autres
Publié: (2025)
par: Sudro, Protima Nomo, et autres
Publié: (2025)
Adaptive Variational Inference in Probabilistic Graphical Models: Beyond Bethe, Tree-Reweighted, and Convex Free Energies
par: Leisenberger, Harald, et autres
Publié: (2025)
par: Leisenberger, Harald, et autres
Publié: (2025)
Documents similaires
-
A multimodal Bayesian Network for symptom-level depression and anxiety prediction from voice and speech data
par: Norbury, Agnes, et autres
Publié: (2025) -
CognoSpeak: an automatic, remote assessment of early cognitive decline in real-world conversational speech
par: Pahar, Madhurananda, et autres
Publié: (2025) -
Online speaker diarization of meetings guided by speech separation
par: Gruttadauria, Elio, et autres
Publié: (2024) -
voice2mode: Phonation Mode Classification in Singing using Self-Supervised Speech Models
par: Justus, Aju Ani, et autres
Publié: (2026) -
Acoustic and perceptual differences between standard and accented speech and their voice clones
par: Yang, Tianle, et autres
Publié: (2026)