Gespeichert in:
| Hauptverfasser: | Ueda, Lucas H., Marques, Leonardo B. de M. M., Simões, Flávio O., Neto, Mário U., Runstein, Fernando, Bó, Bianca Dal, Costa, Paula D. P. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2409.17364 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SelfTTS: cross-speaker style transfer through explicit embedding disentanglement and self-refinement using self-augmentation
von: Ueda, Lucas H., et al.
Veröffentlicht: (2026)
von: Ueda, Lucas H., et al.
Veröffentlicht: (2026)
SponTTS: modeling and transferring spontaneous style for TTS
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
Improving Data Augmentation-based Cross-Speaker Style Transfer for TTS with Singing Voice, Style Filtering, and F0 Matching
von: Marques, Leonardo B. de M. M., et al.
Veröffentlicht: (2024)
von: Marques, Leonardo B. de M. M., et al.
Veröffentlicht: (2024)
Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model
von: Ueda, Lucas, et al.
Veröffentlicht: (2025)
von: Ueda, Lucas, et al.
Veröffentlicht: (2025)
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
Improving curriculum learning for target speaker extraction with synthetic speakers
von: Liu, Yun, et al.
Veröffentlicht: (2024)
von: Liu, Yun, et al.
Veröffentlicht: (2024)
Gender-ambiguous voice generation through feminine speaking style transfer in male voices
von: Koutsogiannaki, Maria, et al.
Veröffentlicht: (2024)
von: Koutsogiannaki, Maria, et al.
Veröffentlicht: (2024)
Hierarchical speaker representation for target speaker extraction
von: He, Shulin, et al.
Veröffentlicht: (2022)
von: He, Shulin, et al.
Veröffentlicht: (2022)
Improved symbolic drum style classification with grammar-based hierarchical representations
von: Géré, Léo, et al.
Veröffentlicht: (2024)
von: Géré, Léo, et al.
Veröffentlicht: (2024)
Combining audio control and style transfer using latent diffusion
von: Demerlé, Nils, et al.
Veröffentlicht: (2024)
von: Demerlé, Nils, et al.
Veröffentlicht: (2024)
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
von: Kunešová, Marie, et al.
Veröffentlicht: (2025)
von: Kunešová, Marie, et al.
Veröffentlicht: (2025)
Scaling NVIDIA's Multi-speaker Multi-lingual TTS Systems with Zero-Shot TTS to Indic Languages
von: Arora, Akshit, et al.
Veröffentlicht: (2024)
von: Arora, Akshit, et al.
Veröffentlicht: (2024)
PromptASR for contextualized ASR with controllable style
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
Crab: Multi Layer Contrastive Supervision to Improve Speech Emotion Recognition Under Both Acted and Natural Speech Condition
von: Ueda, Lucas H., et al.
Veröffentlicht: (2026)
von: Ueda, Lucas H., et al.
Veröffentlicht: (2026)
Accent-VITS:accent transfer for end-to-end TTS
von: Ma, Linhan, et al.
Veröffentlicht: (2023)
von: Ma, Linhan, et al.
Veröffentlicht: (2023)
Improving speaker verification robustness with synthetic emotional utterances
von: Koditala, Nikhil Kumar, et al.
Veröffentlicht: (2024)
von: Koditala, Nikhil Kumar, et al.
Veröffentlicht: (2024)
Complexity of frequency fluctuations and the interpretive style in the bass viola da gamba
von: Lugo, Igor, et al.
Veröffentlicht: (2025)
von: Lugo, Igor, et al.
Veröffentlicht: (2025)
SynthCloner: Synthesizer-style Audio Transfer via Factorized Codec with ADSR Envelope Control
von: Liu, Jeng-Yue, et al.
Veröffentlicht: (2025)
von: Liu, Jeng-Yue, et al.
Veröffentlicht: (2025)
A Dataset for Automatic Assessment of TTS Quality in Spanish
von: Welford, Alejandro Sosa, et al.
Veröffentlicht: (2025)
von: Welford, Alejandro Sosa, et al.
Veröffentlicht: (2025)
Text adaptation for speaker verification with speaker-text factorized embeddings
von: Yang, Yexin, et al.
Veröffentlicht: (2025)
von: Yang, Yexin, et al.
Veröffentlicht: (2025)
X-CrossNet: A complex spectral mapping approach to target speaker extraction with cross attention speaker embedding fusion
von: Sun, Chang, et al.
Veröffentlicht: (2024)
von: Sun, Chang, et al.
Veröffentlicht: (2024)
On the influence of language similarity in non-target speaker verification trials
von: Reuter, Paul M., et al.
Veröffentlicht: (2025)
von: Reuter, Paul M., et al.
Veröffentlicht: (2025)
E1 TTS: Simple and Fast Non-Autoregressive TTS
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
von: Pascu, Octavian, et al.
Veröffentlicht: (2023)
von: Pascu, Octavian, et al.
Veröffentlicht: (2023)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages
von: Qharabagh, Mahta Fetrat, et al.
Veröffentlicht: (2024)
von: Qharabagh, Mahta Fetrat, et al.
Veröffentlicht: (2024)
How phonemes contribute to deep speaker models?
von: Li, Pengqi, et al.
Veröffentlicht: (2024)
von: Li, Pengqi, et al.
Veröffentlicht: (2024)
Word-wise intonation model for cross-language TTS systems
von: A., Tomilov A., et al.
Veröffentlicht: (2024)
von: A., Tomilov A., et al.
Veröffentlicht: (2024)
The importance of spatial and spectral information in multiple speaker tracking
von: Beit-On, Hanan, et al.
Veröffentlicht: (2024)
von: Beit-On, Hanan, et al.
Veröffentlicht: (2024)
Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
von: He, Xinlu, et al.
Veröffentlicht: (2025)
von: He, Xinlu, et al.
Veröffentlicht: (2025)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
von: Xue, Heyang, et al.
Veröffentlicht: (2025)
von: Xue, Heyang, et al.
Veröffentlicht: (2025)
Audio-visual child-adult speaker classification in dyadic interactions
von: Xu, Anfeng, et al.
Veröffentlicht: (2023)
von: Xu, Anfeng, et al.
Veröffentlicht: (2023)
Spoken language change detection inspired by speaker change detection
von: Mishra, Jagabandhu, et al.
Veröffentlicht: (2023)
von: Mishra, Jagabandhu, et al.
Veröffentlicht: (2023)
Transfer the linguistic representations from TTS to accent conversion with non-parallel data
von: Chen, Xi, et al.
Veröffentlicht: (2024)
von: Chen, Xi, et al.
Veröffentlicht: (2024)
Gradient weighting for speaker verification in extremely low Signal-to-Noise Ratio
von: Ma, Yi, et al.
Veröffentlicht: (2024)
von: Ma, Yi, et al.
Veröffentlicht: (2024)
Spectral or spatial? Leveraging both for speaker extraction in challenging data conditions
von: Eisenberg, Aviad, et al.
Veröffentlicht: (2025)
von: Eisenberg, Aviad, et al.
Veröffentlicht: (2025)
Why disentanglement-based speaker anonymization systems fail at preserving emotions?
von: Gaznepoglu, Ünal Ege, et al.
Veröffentlicht: (2025)
von: Gaznepoglu, Ünal Ege, et al.
Veröffentlicht: (2025)
Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
von: Murata, Masato, et al.
Veröffentlicht: (2025)
von: Murata, Masato, et al.
Veröffentlicht: (2025)
HeightCeleb - an enrichment of VoxCeleb dataset with speaker height information
von: Kacprzak, Stanisław, et al.
Veröffentlicht: (2024)
von: Kacprzak, Stanisław, et al.
Veröffentlicht: (2024)
Improving fairness in speaker verification via Group-adapted Fusion Network
von: Shen, Hua, et al.
Veröffentlicht: (2022)
von: Shen, Hua, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
SelfTTS: cross-speaker style transfer through explicit embedding disentanglement and self-refinement using self-augmentation
von: Ueda, Lucas H., et al.
Veröffentlicht: (2026) -
SponTTS: modeling and transferring spontaneous style for TTS
von: Li, Hanzhao, et al.
Veröffentlicht: (2023) -
Improving Data Augmentation-based Cross-Speaker Style Transfer for TTS with Singing Voice, Style Filtering, and F0 Matching
von: Marques, Leonardo B. de M. M., et al.
Veröffentlicht: (2024) -
Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model
von: Ueda, Lucas, et al.
Veröffentlicht: (2025) -
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)