DarkStream: real-time speech anonymization with low latency
Fuente:
arXiv
Guardado en:
| Autores principales: | Quamer, Waris, Gutierrez-Osuna, Ricardo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
End-to-end streaming model for low-latency speech anonymization
por: Quamer, Waris, et al.
Publicado: (2024)
por: Quamer, Waris, et al.
Publicado: (2024)
Disentangling segmental and prosodic factors to non-native speech comprehensibility
por: Quamer, Waris, et al.
Publicado: (2024)
por: Quamer, Waris, et al.
Publicado: (2024)
PHONOS: PHOnetic Neutralization for Online Streaming Applications
por: Quamer, Waris, et al.
Publicado: (2026)
por: Quamer, Waris, et al.
Publicado: (2026)
TVTSyn: Content-Synchronous Time-Varying Timbre for Streaming Voice Conversion and Anonymization
por: Quamer, Waris, et al.
Publicado: (2026)
por: Quamer, Waris, et al.
Publicado: (2026)
A low latency attention module for streaming self-supervised speech representation learning
por: Ma, Jianbo, et al.
Publicado: (2023)
por: Ma, Jianbo, et al.
Publicado: (2023)
Perceptual implications of automatic anonymization in pathological speech
por: Arasteh, Soroosh Tayebi, et al.
Publicado: (2025)
por: Arasteh, Soroosh Tayebi, et al.
Publicado: (2025)
Introduction to speech recognition
por: Dauphin, Gabriel
Publicado: (2024)
por: Dauphin, Gabriel
Publicado: (2024)
Moshi: a speech-text foundation model for real-time dialogue
por: Défossez, Alexandre, et al.
Publicado: (2024)
por: Défossez, Alexandre, et al.
Publicado: (2024)
The evaluation of a code-switched Sepedi-English automatic speech recognition system
por: Phaladi, Amanda, et al.
Publicado: (2024)
por: Phaladi, Amanda, et al.
Publicado: (2024)
Speech foundation models in healthcare: Effect of layer selection on pathological speech feature prediction
por: Wiepert, Daniela A., et al.
Publicado: (2024)
por: Wiepert, Daniela A., et al.
Publicado: (2024)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
por: Fujita, Kenichi, et al.
Publicado: (2024)
por: Fujita, Kenichi, et al.
Publicado: (2024)
Knowledge boosting during low-latency inference
por: Srinivas, Vidya, et al.
Publicado: (2024)
por: Srinivas, Vidya, et al.
Publicado: (2024)
Self-supervised learning of speech representations with Dutch archival data
por: Vaessen, Nik, et al.
Publicado: (2025)
por: Vaessen, Nik, et al.
Publicado: (2025)
Navigating the Minefield of MT Beam Search in Cascaded Streaming Speech Translation
por: Rabatin, Rastislav, et al.
Publicado: (2024)
por: Rabatin, Rastislav, et al.
Publicado: (2024)
SpeakStream: Streaming Text-to-Speech with Interleaved Data
por: Bai, Richard He, et al.
Publicado: (2025)
por: Bai, Richard He, et al.
Publicado: (2025)
Non-verbal information in spontaneous speech -- towards a new framework of analysis
por: Biron, Tirza, et al.
Publicado: (2024)
por: Biron, Tirza, et al.
Publicado: (2024)
Deferred NAM: Low-latency Top-K Context Injection via Deferred Context Encoding for Non-Streaming ASR
por: Wu, Zelin, et al.
Publicado: (2024)
por: Wu, Zelin, et al.
Publicado: (2024)
WhisperRT -- Turning Whisper into a Causal Streaming Model
por: Krichli, Tomer, et al.
Publicado: (2025)
por: Krichli, Tomer, et al.
Publicado: (2025)
Efficient Adapter Finetuning for Tail Languages in Streaming Multilingual ASR
por: Bai, Junwen, et al.
Publicado: (2024)
por: Bai, Junwen, et al.
Publicado: (2024)
IIITH-BUT system for IWSLT 2025 low-resource Bhojpuri to Hindi speech translation
por: Akkiraju, Bhavana, et al.
Publicado: (2025)
por: Akkiraju, Bhavana, et al.
Publicado: (2025)
A multilingual training strategy for low resource Text to Speech
por: Amalas, Asma, et al.
Publicado: (2024)
por: Amalas, Asma, et al.
Publicado: (2024)
Prominence-aware automatic speech recognition for conversational speech
por: Linke, Julian, et al.
Publicado: (2025)
por: Linke, Julian, et al.
Publicado: (2025)
Predicting positive transfer for improved low-resource speech recognition using acoustic pseudo-tokens
por: San, Nay, et al.
Publicado: (2024)
por: San, Nay, et al.
Publicado: (2024)
Two-component spatiotemporal template for activation-inhibition of speech in ECoG
por: Easthope, Eric
Publicado: (2024)
por: Easthope, Eric
Publicado: (2024)
Technical Report on classification of literature related to children speech disorder
por: Wang, Ziang, et al.
Publicado: (2025)
por: Wang, Ziang, et al.
Publicado: (2025)
More than words: Advancements and challenges in speech recognition for singing
por: Kruspe, Anna
Publicado: (2024)
por: Kruspe, Anna
Publicado: (2024)
End to end Hindi to English speech conversion using Bark, mBART and a finetuned XLSR Wav2Vec2
por: Tathe, Aniket, et al.
Publicado: (2024)
por: Tathe, Aniket, et al.
Publicado: (2024)
Developing multilingual speech synthesis system for Ojibwe, Mi'kmaq, and Maliseet
por: Wang, Shenran, et al.
Publicado: (2025)
por: Wang, Shenran, et al.
Publicado: (2025)
Target speaker anonymization in multi-speaker recordings
por: Tomashenko, Natalia, et al.
Publicado: (2025)
por: Tomashenko, Natalia, et al.
Publicado: (2025)
"KAN you hear me?" Exploring Kolmogorov-Arnold Networks for Spoken Language Understanding
por: Koudounas, Alkis, et al.
Publicado: (2025)
por: Koudounas, Alkis, et al.
Publicado: (2025)
Robust Unsupervised Adaptation of a Speech Recogniser Using Entropy Minimisation and Speaker Codes
por: van Dalen, Rogier C., et al.
Publicado: (2025)
por: van Dalen, Rogier C., et al.
Publicado: (2025)
Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses
por: Kim, Sungnyun, et al.
Publicado: (2025)
por: Kim, Sungnyun, et al.
Publicado: (2025)
Mitigating Data Imbalance in Automated Speaking Assessment
por: Tsai, Fong-Chun, et al.
Publicado: (2025)
por: Tsai, Fong-Chun, et al.
Publicado: (2025)
Re-evaluating Minimum Bayes Risk Decoding for Automatic Speech Recognition
por: Jinnai, Yuu
Publicado: (2025)
por: Jinnai, Yuu
Publicado: (2025)
Scalable Frameworks for Real-World Audio-Visual Speech Recognition
por: Kim, Sungnyun
Publicado: (2025)
por: Kim, Sungnyun
Publicado: (2025)
Data-Centric Lessons To Improve Speech-Language Pretraining
por: Udandarao, Vishaal, et al.
Publicado: (2025)
por: Udandarao, Vishaal, et al.
Publicado: (2025)
DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation
por: Song, Yakun, et al.
Publicado: (2025)
por: Song, Yakun, et al.
Publicado: (2025)
Gender Bias in Instruction-Guided Speech Synthesis Models
por: Kuan, Chun-Yi, et al.
Publicado: (2025)
por: Kuan, Chun-Yi, et al.
Publicado: (2025)
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition
por: Kim, Sungnyun, et al.
Publicado: (2025)
por: Kim, Sungnyun, et al.
Publicado: (2025)
Spoken Language Understanding on Unseen Tasks With In-Context Learning
por: Agrawal, Neeraj, et al.
Publicado: (2025)
por: Agrawal, Neeraj, et al.
Publicado: (2025)
Ejemplares similares
-
End-to-end streaming model for low-latency speech anonymization
por: Quamer, Waris, et al.
Publicado: (2024) -
Disentangling segmental and prosodic factors to non-native speech comprehensibility
por: Quamer, Waris, et al.
Publicado: (2024) -
PHONOS: PHOnetic Neutralization for Online Streaming Applications
por: Quamer, Waris, et al.
Publicado: (2026) -
TVTSyn: Content-Synchronous Time-Varying Timbre for Streaming Voice Conversion and Anonymization
por: Quamer, Waris, et al.
Publicado: (2026) -
A low latency attention module for streaming self-supervised speech representation learning
por: Ma, Jianbo, et al.
Publicado: (2023)