Voxceleb-ESP: preliminary experiments detecting Spanish celebrities from their voices
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Labrador, Beltrán, Otero-Gonzalez, Manuel, Lozano-Diez, Alicia, Ramos, Daniel, Toledano, Doroteo T., Gonzalez-Rodriguez, Joaquin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Speaker Embeddings in End-to-End Neural Diarization for Two-Speaker Scenarios
von: Alvarez-Trejos, Juan Ignacio, et al.
Veröffentlicht: (2024)
von: Alvarez-Trejos, Juan Ignacio, et al.
Veröffentlicht: (2024)
Towards detecting the pathological subharmonic voicing with fully convolutional neural networks
von: Ikuma, Takeshi, et al.
Veröffentlicht: (2025)
von: Ikuma, Takeshi, et al.
Veröffentlicht: (2025)
Introducing voice timbre attribute detection
von: He, Jinghao, et al.
Veröffentlicht: (2025)
von: He, Jinghao, et al.
Veröffentlicht: (2025)
Gender-ambiguous voice generation through feminine speaking style transfer in male voices
von: Koutsogiannaki, Maria, et al.
Veröffentlicht: (2024)
von: Koutsogiannaki, Maria, et al.
Veröffentlicht: (2024)
Easy, Interpretable, Effective: openSMILE for voice deepfake detection
von: Pascu, Octavian, et al.
Veröffentlicht: (2024)
von: Pascu, Octavian, et al.
Veröffentlicht: (2024)
Comparison of fundamental frequency estimators with subharmonic voice signals
von: Ikuma, Takeshi, et al.
Veröffentlicht: (2025)
von: Ikuma, Takeshi, et al.
Veröffentlicht: (2025)
Subjective quality evaluation of personalized own voice reconstruction systems
von: Ohlenbusch, Mattes, et al.
Veröffentlicht: (2025)
von: Ohlenbusch, Mattes, et al.
Veröffentlicht: (2025)
DNN-based ensemble singing voice synthesis with interactions between singers
von: Hyodo, Hiroaki, et al.
Veröffentlicht: (2024)
von: Hyodo, Hiroaki, et al.
Veröffentlicht: (2024)
Using voice analysis as an early indicator of risk for depression in young adults
von: Scherer, Klaus R., et al.
Veröffentlicht: (2024)
von: Scherer, Klaus R., et al.
Veröffentlicht: (2024)
Adversarial speech for voice privacy protection from Personalized Speech generation
von: Chen, Shihao, et al.
Veröffentlicht: (2024)
von: Chen, Shihao, et al.
Veröffentlicht: (2024)
Building speech corpus with diverse voice characteristics for its prompt-based representation
von: Watanabe, Aya, et al.
Veröffentlicht: (2024)
von: Watanabe, Aya, et al.
Veröffentlicht: (2024)
Audiovisual angle and voice incongruence do not affect audiovisual verbal short-term memory in virtual reality
von: Ermert, Cosima A., et al.
Veröffentlicht: (2024)
von: Ermert, Cosima A., et al.
Veröffentlicht: (2024)
A Dataset for Automatic Assessment of TTS Quality in Spanish
von: Welford, Alejandro Sosa, et al.
Veröffentlicht: (2025)
von: Welford, Alejandro Sosa, et al.
Veröffentlicht: (2025)
Accurate analysis of the pitch pulse-based magnitude/phase structure of natural vowels and assessment of three lightweight time/frequency voicing restoration methods
von: Ferreira, Aníbal J. S., et al.
Veröffentlicht: (2025)
von: Ferreira, Aníbal J. S., et al.
Veröffentlicht: (2025)
Spoken language change detection inspired by speaker change detection
von: Mishra, Jagabandhu, et al.
Veröffentlicht: (2023)
von: Mishra, Jagabandhu, et al.
Veröffentlicht: (2023)
Absorbing Discrete Diffusion for Speech Enhancement
von: Gonzalez, Philippe
Veröffentlicht: (2026)
von: Gonzalez, Philippe
Veröffentlicht: (2026)
Sound event localization and detection based on crnn using rectangular filters and channel rotation data augmentation
von: Ronchini, Francesca, et al.
Veröffentlicht: (2020)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2020)
Resource-constrained stereo singing voice cancellation
von: Borrelli, Clara, et al.
Veröffentlicht: (2024)
von: Borrelli, Clara, et al.
Veröffentlicht: (2024)
DiaPer: End-to-End Neural Diarization with Perceiver-Based Attractors
von: Landini, Federico, et al.
Veröffentlicht: (2023)
von: Landini, Federico, et al.
Veröffentlicht: (2023)
Are audio DeepFake detection models polyglots?
von: Marek, Bartłomiej, et al.
Veröffentlicht: (2024)
von: Marek, Bartłomiej, et al.
Veröffentlicht: (2024)
Frequency-aware convolution for sound event detection
von: Song, Tao, et al.
Veröffentlicht: (2024)
von: Song, Tao, et al.
Veröffentlicht: (2024)
Onset and offset weighted loss function for sound event detection
von: Song, Tao
Veröffentlicht: (2024)
von: Song, Tao
Veröffentlicht: (2024)
Fine-tune the pretrained ATST model for sound event detection
von: Shao, Nian, et al.
Veröffentlicht: (2023)
von: Shao, Nian, et al.
Veröffentlicht: (2023)
Controllable joint noise reduction and hearing loss compensation using a differentiable auditory model
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2025)
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2025)
Do End-to-End Neural Diarization Attractors Need to Encode Speaker Characteristic Information?
von: Zhang, Lin, et al.
Veröffentlicht: (2024)
von: Zhang, Lin, et al.
Veröffentlicht: (2024)
Leveraging Self-Supervised Learning for Speaker Diarization
von: Han, Jiangyu, et al.
Veröffentlicht: (2024)
von: Han, Jiangyu, et al.
Veröffentlicht: (2024)
Non-autoregressive real-time Accent Conversion model with voice cloning
von: Nechaev, Vladimir, et al.
Veröffentlicht: (2024)
von: Nechaev, Vladimir, et al.
Veröffentlicht: (2024)
Representational learning for an anomalous sound detection system with source separation model
von: Shin, Seunghyeon, et al.
Veröffentlicht: (2024)
von: Shin, Seunghyeon, et al.
Veröffentlicht: (2024)
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
von: Pascu, Octavian, et al.
Veröffentlicht: (2023)
von: Pascu, Octavian, et al.
Veröffentlicht: (2023)
A robust audio deepfake detection system via multi-view feature
von: Yang, Yujie, et al.
Veröffentlicht: (2024)
von: Yang, Yujie, et al.
Veröffentlicht: (2024)
Joint Training of Speaker Embedding Extractor, Speech and Overlap Detection for Diarization
von: Pálka, Petr, et al.
Veröffentlicht: (2024)
von: Pálka, Petr, et al.
Veröffentlicht: (2024)
Full-frequency dynamic convolution: a physical frequency-dependent convolution for sound event detection
von: Yue, Haobo, et al.
Veröffentlicht: (2024)
von: Yue, Haobo, et al.
Veröffentlicht: (2024)
Stereo sound event localization and detection based on PSELDnet pretraining and BiMamba sequence modeling
von: Gao, Wenmiao, et al.
Veröffentlicht: (2025)
von: Gao, Wenmiao, et al.
Veröffentlicht: (2025)
Performance and energy balance: a comprehensive study of state-of-the-art sound event detection systems
von: Ronchini, Francesca, et al.
Veröffentlicht: (2023)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2023)
Automatic acoustic detection of birds through deep learning: the first Bird Audio Detection challenge
von: Stowell, Dan, et al.
Veröffentlicht: (2018)
von: Stowell, Dan, et al.
Veröffentlicht: (2018)
PAGURI: a user experience study of creative interaction with text-to-music models
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2024)
End-to-End Multi-Task Learning for Adjustable Joint Noise Reduction and Hearing Loss Compensation
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2026)
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2026)
Can we reconstruct a dysarthric voice with the large speech model Parler TTS?
von: Sanchez, Ariadna, et al.
Veröffentlicht: (2025)
von: Sanchez, Ariadna, et al.
Veröffentlicht: (2025)
Resnet-conformer network with shared weights and attention mechanism for sound event localization, detection, and distance estimation
von: Vo, Quoc Thinh, et al.
Veröffentlicht: (2025)
von: Vo, Quoc Thinh, et al.
Veröffentlicht: (2025)
Sound event detection based on auxiliary decoder and maximum probability aggregation for DCASE Challenge 2024 Task 4
von: Son, Sang Won, et al.
Veröffentlicht: (2024)
von: Son, Sang Won, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Leveraging Speaker Embeddings in End-to-End Neural Diarization for Two-Speaker Scenarios
von: Alvarez-Trejos, Juan Ignacio, et al.
Veröffentlicht: (2024) -
Towards detecting the pathological subharmonic voicing with fully convolutional neural networks
von: Ikuma, Takeshi, et al.
Veröffentlicht: (2025) -
Introducing voice timbre attribute detection
von: He, Jinghao, et al.
Veröffentlicht: (2025) -
Gender-ambiguous voice generation through feminine speaking style transfer in male voices
von: Koutsogiannaki, Maria, et al.
Veröffentlicht: (2024) -
Easy, Interpretable, Effective: openSMILE for voice deepfake detection
von: Pascu, Octavian, et al.
Veröffentlicht: (2024)