DINO-VITS: Data-Efficient Zero-Shot TTS with Self-Supervised Speaker Verification Loss for Noise Robustness
Fuente:
arXiv
Saved in:
| Main Authors: | Pankov, Vikentii, Pronina, Valeria, Kuzmin, Alexander, Borisov, Maksim, Usoltsev, Nikita, Zeng, Xingshan, Golubkov, Alexander, Ermolenko, Nikolai, Shirshova, Aleksandra, Matveeva, Yulia |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PFluxTTS: Hybrid Flow-Matching TTS with Robust Cross-Lingual Voice Cloning and Inference-Time Model Fusion
by: Pankov, Vikentii, et al.
Published: (2026)
by: Pankov, Vikentii, et al.
Published: (2026)
Accent-VITS:accent transfer for end-to-end TTS
by: Ma, Linhan, et al.
Published: (2023)
by: Ma, Linhan, et al.
Published: (2023)
Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness
by: Feng, Xincan, et al.
Published: (2024)
by: Feng, Xincan, et al.
Published: (2024)
NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech
by: Borisov, Maksim, et al.
Published: (2025)
by: Borisov, Maksim, et al.
Published: (2025)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
by: Jeon, Yejin, et al.
Published: (2024)
by: Jeon, Yejin, et al.
Published: (2024)
Monomial stability of Frobenius images
by: Borisov, Nikita
Published: (2025)
by: Borisov, Nikita
Published: (2025)
Comparative Evaluation of Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2
by: Rackauckas, Zackary, et al.
Published: (2025)
by: Rackauckas, Zackary, et al.
Published: (2025)
BnTTS: Few-Shot Speaker Adaptation in Low-Resource Setting
by: Basher, Mohammad Jahid Ibna, et al.
Published: (2025)
by: Basher, Mohammad Jahid Ibna, et al.
Published: (2025)
Locally Integer Polynomial Functions
by: Borisov, Alexander
Published: (2024)
by: Borisov, Alexander
Published: (2024)
A Structure Sheaf for Kirch Topology
by: Borisov, Alexander
Published: (2026)
by: Borisov, Alexander
Published: (2026)
Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS
by: Wang, Haoyu, et al.
Published: (2024)
by: Wang, Haoyu, et al.
Published: (2024)
StreamVoiceAnon+: Emotion-Preserving Streaming Speaker Anonymization via Frame-Level Acoustic Distillation
by: Kuzmin, Nikita, et al.
Published: (2026)
by: Kuzmin, Nikita, et al.
Published: (2026)
Motives, cohomological invariants and Freudenthal magic square
by: Geldhauser, Nikita, et al.
Published: (2026)
by: Geldhauser, Nikita, et al.
Published: (2026)
An Investigation of Noise Robustness for Flow-Matching-Based Zero-Shot TTS
by: Wang, Xiaofei, et al.
Published: (2024)
by: Wang, Xiaofei, et al.
Published: (2024)
Time-Layer Adaptive Alignment for Speaker Similarity in Flow-Matching Based Zero-Shot TTS
by: Li, Haoyu, et al.
Published: (2025)
by: Li, Haoyu, et al.
Published: (2025)
Zero-Shot vs. Few-Shot Multi-Speaker TTS Using Pre-trained Czech SpeechT5 Model
by: Lehečka, Jan, et al.
Published: (2024)
by: Lehečka, Jan, et al.
Published: (2024)
VITS : Variational Inference Thompson Sampling for contextual bandits
by: Clavier, Pierre, et al.
Published: (2023)
by: Clavier, Pierre, et al.
Published: (2023)
Porosity and topological properties of triply periodic minimal surfaces
by: Ermolenko, Sergei, et al.
Published: (2024)
by: Ermolenko, Sergei, et al.
Published: (2024)
On the integrality of some P-recursive sequences
by: Matveeva, Anastasia
Published: (2025)
by: Matveeva, Anastasia
Published: (2025)
Noise-Conditioned Mixture-of-Experts Framework for Robust Speaker Verification
by: Gu, Bin, et al.
Published: (2025)
by: Gu, Bin, et al.
Published: (2025)
Zero-Shot Multi-Lingual Speaker Verification in Clinical Trials
by: Akram, Ali, et al.
Published: (2024)
by: Akram, Ali, et al.
Published: (2024)
Stream-Voice-Anon: Enhancing Utility of Real-Time Speaker Anonymization via Neural Audio Codec and Language Models
by: Kuzmin, Nikita, et al.
Published: (2026)
by: Kuzmin, Nikita, et al.
Published: (2026)
Low-resource Machine Translation for Code-switched Kazakh-Russian Language Pair
by: Borisov, Maksim, et al.
Published: (2025)
by: Borisov, Maksim, et al.
Published: (2025)
Energy-conserving intermittent-contact motion in complex models
by: Pankov, Sergey
Published: (2024)
by: Pankov, Sergey
Published: (2024)
Symmetric $(2^k-1,2^{k-1},2^{k-2})$-designs which are $(2^{k-1}-1)$-pyramidal over abelian groups
by: Pankov, Mark
Published: (2025)
by: Pankov, Mark
Published: (2025)
A non-surjective Wigner-type theorem in terms of equivalent pairs of subspaces
by: Pankov, Mark
Published: (2024)
by: Pankov, Mark
Published: (2024)
Recovering of the Grassmann graph from the subgraph of non-degenerate subspaces
by: Pankov, Mark
Published: (2026)
by: Pankov, Mark
Published: (2026)
Asymptotics of immaculate line bundles on smooth toric Deligne-Mumford stacks
by: Borisov, Lev, et al.
Published: (2023)
by: Borisov, Lev, et al.
Published: (2023)
IsoChronoMeter: A simple and effective isochronic translation evaluation metric
by: Rozanov, Nikolai, et al.
Published: (2024)
by: Rozanov, Nikolai, et al.
Published: (2024)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
by: Fu, Ruibo, et al.
Published: (2024)
by: Fu, Ruibo, et al.
Published: (2024)
Few-Shot Adaptation of Grounding DINO for Agricultural Domain
by: Singh, Rajhans, et al.
Published: (2025)
by: Singh, Rajhans, et al.
Published: (2025)
Z3D: Zero-Shot 3D Visual Grounding from Images
by: Drozdov, Nikita, et al.
Published: (2026)
by: Drozdov, Nikita, et al.
Published: (2026)
A Joint Noise Disentanglement and Adversarial Training Framework for Robust Speaker Verification
by: Xing, Xujiang, et al.
Published: (2024)
by: Xing, Xujiang, et al.
Published: (2024)
Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
by: He, Xinlu, et al.
Published: (2025)
by: He, Xinlu, et al.
Published: (2025)
PAVITS: Exploring Prosody-aware VITS for End-to-End Emotional Voice Conversion
by: Qi, Tianhua, et al.
Published: (2024)
by: Qi, Tianhua, et al.
Published: (2024)
ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
by: Kim, Minu, et al.
Published: (2025)
by: Kim, Minu, et al.
Published: (2025)
DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation
by: Meng, Ming, et al.
Published: (2025)
by: Meng, Ming, et al.
Published: (2025)
DINO-Foresight: Looking into the Future with DINO
by: Karypidis, Efstathios, et al.
Published: (2024)
by: Karypidis, Efstathios, et al.
Published: (2024)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
by: Eskimez, Sefik Emre, et al.
Published: (2024)
by: Eskimez, Sefik Emre, et al.
Published: (2024)
TTS-1 Technical Report
by: Atamanenko, Oleg, et al.
Published: (2025)
by: Atamanenko, Oleg, et al.
Published: (2025)
Similar Items
-
PFluxTTS: Hybrid Flow-Matching TTS with Robust Cross-Lingual Voice Cloning and Inference-Time Model Fusion
by: Pankov, Vikentii, et al.
Published: (2026) -
Accent-VITS:accent transfer for end-to-end TTS
by: Ma, Linhan, et al.
Published: (2023) -
Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness
by: Feng, Xincan, et al.
Published: (2024) -
NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech
by: Borisov, Maksim, et al.
Published: (2025) -
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
by: Jeon, Yejin, et al.
Published: (2024)