Removing Speaker Information from Speech Representation using Variable-Length Soft Pooling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hwang, Injune, Lee, Kyogu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
von: Han, Seungu, et al.
Veröffentlicht: (2026)
von: Han, Seungu, et al.
Veröffentlicht: (2026)
Towards Bitrate-Efficient and Noise-Robust Speech Coding with Variable Bitrate RVQ
von: Chae, Yunkee, et al.
Veröffentlicht: (2025)
von: Chae, Yunkee, et al.
Veröffentlicht: (2025)
Hear Your Face: Face-based voice conversion with F0 estimation
von: Lee, Jaejun, et al.
Veröffentlicht: (2024)
von: Lee, Jaejun, et al.
Veröffentlicht: (2024)
Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation
von: Lee, Jaejun, et al.
Veröffentlicht: (2025)
von: Lee, Jaejun, et al.
Veröffentlicht: (2025)
Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2025)
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2025)
Music Auto-Tagging with Robust Music Representation Learned via Domain Adversarial Training
von: Joung, Haesun, et al.
Veröffentlicht: (2024)
von: Joung, Haesun, et al.
Veröffentlicht: (2024)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
von: Shi, Hao, et al.
Veröffentlicht: (2024)
von: Shi, Hao, et al.
Veröffentlicht: (2024)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
Practical and Reproducible Symbolic Music Generation by Large Language Models with Structural Embeddings
von: Rhyu, Seungyeon, et al.
Veröffentlicht: (2024)
von: Rhyu, Seungyeon, et al.
Veröffentlicht: (2024)
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
Perceiver-Prompt: Flexible Speaker Adaptation in Whisper for Chinese Disordered Speech Recognition
von: Jiang, Yicong, et al.
Veröffentlicht: (2024)
von: Jiang, Yicong, et al.
Veröffentlicht: (2024)
Enhancing Speech Emotion Recognition through Segmental Average Pooling of Self-Supervised Learning Features
von: Hyeon, Jonghwan, et al.
Veröffentlicht: (2024)
von: Hyeon, Jonghwan, et al.
Veröffentlicht: (2024)
Differentiable Modal Synthesis for Physical Modeling of Planar String Sound and Motion Simulation
von: Lee, Jin Woo, et al.
Veröffentlicht: (2024)
von: Lee, Jin Woo, et al.
Veröffentlicht: (2024)
Few-step Adversarial Schrödinger Bridge for Generative Speech Enhancement
von: Han, Seungu, et al.
Veröffentlicht: (2025)
von: Han, Seungu, et al.
Veröffentlicht: (2025)
Distance Sampling-based Paraphraser Leveraging ChatGPT for Text Data Manipulation
von: Oh, Yoori, et al.
Veröffentlicht: (2024)
von: Oh, Yoori, et al.
Veröffentlicht: (2024)
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
von: Han, Wooseok, et al.
Veröffentlicht: (2024)
von: Han, Wooseok, et al.
Veröffentlicht: (2024)
Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech
von: Kim, Taesoo, et al.
Veröffentlicht: (2025)
von: Kim, Taesoo, et al.
Veröffentlicht: (2025)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2024)
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2024)
Speaker Embeddings to Improve Tracking of Intermittent and Moving Speakers
von: Iatariene, Taous, et al.
Veröffentlicht: (2025)
von: Iatariene, Taous, et al.
Veröffentlicht: (2025)
The Sounds of Home: A Speech-Removed Residential Audio Dataset for Sound Event Detection
von: Bibbó, Gabriel, et al.
Veröffentlicht: (2024)
von: Bibbó, Gabriel, et al.
Veröffentlicht: (2024)
WavShape: Information-Theoretic Speech Representation Learning for Fair and Privacy-Aware Audio Processing
von: Baser, Oguzhan, et al.
Veröffentlicht: (2025)
von: Baser, Oguzhan, et al.
Veröffentlicht: (2025)
Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization
von: Bhuiyan, Mohammed Aman, et al.
Veröffentlicht: (2026)
von: Bhuiyan, Mohammed Aman, et al.
Veröffentlicht: (2026)
Memory-Efficient Training for Deep Speaker Embedding Learning in Speaker Verification
von: Liu, Bei, et al.
Veröffentlicht: (2024)
von: Liu, Bei, et al.
Veröffentlicht: (2024)
EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
Asynchronous Voice Anonymization Using Adversarial Perturbation On Speaker Embedding
von: Wang, Rui, et al.
Veröffentlicht: (2024)
von: Wang, Rui, et al.
Veröffentlicht: (2024)
Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
von: Zhang, Hanglei, et al.
Veröffentlicht: (2025)
von: Zhang, Hanglei, et al.
Veröffentlicht: (2025)
Leveraging Speaker Embeddings in End-to-End Neural Diarization for Two-Speaker Scenarios
von: Alvarez-Trejos, Juan Ignacio, et al.
Veröffentlicht: (2024)
von: Alvarez-Trejos, Juan Ignacio, et al.
Veröffentlicht: (2024)
SCRAPS: Speech Contrastive Representations of Acoustic and Phonetic Spaces
von: Vallés-Pérez, Ivan, et al.
Veröffentlicht: (2023)
von: Vallés-Pérez, Ivan, et al.
Veröffentlicht: (2023)
Color-based Emotion Representation for Speech Emotion Recognition
von: Nagase, Ryotaro, et al.
Veröffentlicht: (2026)
von: Nagase, Ryotaro, et al.
Veröffentlicht: (2026)
Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks
von: Lee, Seo-Hyun, et al.
Veröffentlicht: (2023)
von: Lee, Seo-Hyun, et al.
Veröffentlicht: (2023)
Improving Speech Inversion Through Self-Supervised Embeddings and Enhanced Tract Variables
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2023)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2023)
DOSE : Drum One-Shot Extraction from Music Mixture
von: Hwang, Suntae, et al.
Veröffentlicht: (2025)
von: Hwang, Suntae, et al.
Veröffentlicht: (2025)
Text-dependent Speaker Verification (TdSV) Challenge 2024: Challenge Evaluation Plan
von: Hossein, Zeinali, et al.
Veröffentlicht: (2024)
von: Hossein, Zeinali, et al.
Veröffentlicht: (2024)
Disentangling Age and Identity with a Mutual Information Minimization Approach for Cross-Age Speaker Verification
von: Zhang, Fengrun, et al.
Veröffentlicht: (2024)
von: Zhang, Fengrun, et al.
Veröffentlicht: (2024)
Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
Length-Aware Rotary Position Embedding for Text-Speech Alignment
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025)
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025)
Explainable Attribute-Based Speaker Verification
von: Wu, Xiaoliang, et al.
Veröffentlicht: (2024)
von: Wu, Xiaoliang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024) -
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
von: Han, Seungu, et al.
Veröffentlicht: (2026) -
Towards Bitrate-Efficient and Noise-Robust Speech Coding with Variable Bitrate RVQ
von: Chae, Yunkee, et al.
Veröffentlicht: (2025) -
Hear Your Face: Face-based voice conversion with F0 estimation
von: Lee, Jaejun, et al.
Veröffentlicht: (2024) -
Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation
von: Lee, Jaejun, et al.
Veröffentlicht: (2025)