Guardado en:
| Autores principales: | De Silva, Dashanka, Cai, Siqi, Pahuja, Saurav, Schultz, Tanja, Li, Haizhou |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2409.02489 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction
por: Fan, Cunhang, et al.
Publicado: (2025)
por: Fan, Cunhang, et al.
Publicado: (2025)
Low-latency auditory spatial attention detection based on spectro-spatial features from EEG
por: Cai, Siqi, et al.
Publicado: (2021)
por: Cai, Siqi, et al.
Publicado: (2021)
ExPO: Explainable Phonetic Trait-Oriented Network for Speaker Verification
por: Ma, Yi, et al.
Publicado: (2025)
por: Ma, Yi, et al.
Publicado: (2025)
Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction
por: Pan, Zexu, et al.
Publicado: (2025)
por: Pan, Zexu, et al.
Publicado: (2025)
Neuro-MSBG: An End-to-End Neural Model for Hearing Loss Simulation
por: Yuan, Hui-Guan, et al.
Publicado: (2025)
por: Yuan, Hui-Guan, et al.
Publicado: (2025)
Audio-Visual Target Speaker Extraction with Reverse Selective Auditory Attention
por: Tao, Ruijie, et al.
Publicado: (2024)
por: Tao, Ruijie, et al.
Publicado: (2024)
NeuroAMP: A Novel End-to-end General Purpose Deep Neural Amplifier for Personalized Hearing Aids
por: Ahmed, Shafique, et al.
Publicado: (2025)
por: Ahmed, Shafique, et al.
Publicado: (2025)
Multi-Level Speaker Representation for Target Speaker Extraction
por: Zhang, Ke, et al.
Publicado: (2024)
por: Zhang, Ke, et al.
Publicado: (2024)
Target Speaker Extraction through Comparing Noisy Positive and Negative Audio Enrollments
por: Xu, Shitong, et al.
Publicado: (2025)
por: Xu, Shitong, et al.
Publicado: (2025)
NeuroVoz: a Castillian Spanish corpus of parkinsonian speech
por: Mendes-Laureano, Janaína, et al.
Publicado: (2024)
por: Mendes-Laureano, Janaína, et al.
Publicado: (2024)
USED: Universal Speaker Extraction and Diarization
por: Ao, Junyi, et al.
Publicado: (2023)
por: Ao, Junyi, et al.
Publicado: (2023)
Speaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation Algorithm
por: Li, Zhaoyang, et al.
Publicado: (2025)
por: Li, Zhaoyang, et al.
Publicado: (2025)
Speaker Embeddings to Improve Tracking of Intermittent and Moving Speakers
por: Iatariene, Taous, et al.
Publicado: (2025)
por: Iatariene, Taous, et al.
Publicado: (2025)
Breaking Resource Barriers in Speech Emotion Recognition via Data Distillation
por: Chang, Yi, et al.
Publicado: (2024)
por: Chang, Yi, et al.
Publicado: (2024)
On the effectiveness of enrollment speech augmentation for Target Speaker Extraction
por: Li, Junjie, et al.
Publicado: (2024)
por: Li, Junjie, et al.
Publicado: (2024)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
por: Kang, Jiawen, et al.
Publicado: (2024)
por: Kang, Jiawen, et al.
Publicado: (2024)
Memory-Efficient Training for Deep Speaker Embedding Learning in Speaker Verification
por: Liu, Bei, et al.
Publicado: (2024)
por: Liu, Bei, et al.
Publicado: (2024)
Disentangling Age and Identity with a Mutual Information Minimization Approach for Cross-Age Speaker Verification
por: Zhang, Fengrun, et al.
Publicado: (2024)
por: Zhang, Fengrun, et al.
Publicado: (2024)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
por: Kim, Ji-Hoon, et al.
Publicado: (2024)
por: Kim, Ji-Hoon, et al.
Publicado: (2024)
Leveraging Speaker Embeddings in End-to-End Neural Diarization for Two-Speaker Scenarios
por: Alvarez-Trejos, Juan Ignacio, et al.
Publicado: (2024)
por: Alvarez-Trejos, Juan Ignacio, et al.
Publicado: (2024)
Sync-TVA: A Graph-Attention Framework for Multimodal Emotion Recognition with Cross-Modal Fusion
por: Deng, Zeyu, et al.
Publicado: (2025)
por: Deng, Zeyu, et al.
Publicado: (2025)
Investigating Effective Speaker Property Privacy Protection in Federated Learning for Speech Emotion Recognition
por: Tan, Chao, et al.
Publicado: (2024)
por: Tan, Chao, et al.
Publicado: (2024)
Explainable Attribute-Based Speaker Verification
por: Wu, Xiaoliang, et al.
Publicado: (2024)
por: Wu, Xiaoliang, et al.
Publicado: (2024)
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
por: Cho, Deok-Hyeon, et al.
Publicado: (2025)
por: Cho, Deok-Hyeon, et al.
Publicado: (2025)
Affect Decoding in Phonated and Silent Speech Production from Surface EMG
por: Pistrosch, Simon, et al.
Publicado: (2026)
por: Pistrosch, Simon, et al.
Publicado: (2026)
From Modular to End-to-End Speaker Diarization
por: Landini, Federico
Publicado: (2024)
por: Landini, Federico
Publicado: (2024)
Certification of Speaker Recognition Models to Additive Perturbations
por: Korzh, Dmitrii, et al.
Publicado: (2024)
por: Korzh, Dmitrii, et al.
Publicado: (2024)
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction
por: Emon, Jakaria Islam, et al.
Publicado: (2025)
por: Emon, Jakaria Islam, et al.
Publicado: (2025)
MOSA: Music Motion with Semantic Annotation Dataset for Cross-Modal Music Processing
por: Huang, Yu-Fen, et al.
Publicado: (2024)
por: Huang, Yu-Fen, et al.
Publicado: (2024)
ED-sKWS: Early-Decision Spiking Neural Networks for Rapid,and Energy-Efficient Keyword Spotting
por: Song, Zeyang, et al.
Publicado: (2024)
por: Song, Zeyang, et al.
Publicado: (2024)
CrossMuSim: A Cross-Modal Framework for Music Similarity Retrieval with LLM-Powered Text Description Sourcing and Mining
por: Tsoi, Tristan, et al.
Publicado: (2025)
por: Tsoi, Tristan, et al.
Publicado: (2025)
SDBench: A Comprehensive Benchmark Suite for Speaker Diarization
por: Pacheco, Eduardo, et al.
Publicado: (2025)
por: Pacheco, Eduardo, et al.
Publicado: (2025)
The VoxCeleb Speaker Recognition Challenge: A Retrospective
por: Huh, Jaesung, et al.
Publicado: (2024)
por: Huh, Jaesung, et al.
Publicado: (2024)
Listening and Seeing Again: Generative Error Correction for Audio-Visual Speech Recognition
por: Liu, Rui, et al.
Publicado: (2025)
por: Liu, Rui, et al.
Publicado: (2025)
End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization
por: Singh, Prachi, et al.
Publicado: (2024)
por: Singh, Prachi, et al.
Publicado: (2024)
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
por: Guo, Yiwei, et al.
Publicado: (2024)
por: Guo, Yiwei, et al.
Publicado: (2024)
Unispeaker: A Unified Approach for Multimodality-driven Speaker Generation
por: Sheng, Zhengyan, et al.
Publicado: (2025)
por: Sheng, Zhengyan, et al.
Publicado: (2025)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
por: Melechovsky, Jan, et al.
Publicado: (2024)
por: Melechovsky, Jan, et al.
Publicado: (2024)
Asynchronous Voice Anonymization Using Adversarial Perturbation On Speaker Embedding
por: Wang, Rui, et al.
Publicado: (2024)
por: Wang, Rui, et al.
Publicado: (2024)
Evaluating Speaker Identity Coding in Self-supervised Models and Humans
por: Elbanna, Gasser
Publicado: (2024)
por: Elbanna, Gasser
Publicado: (2024)
Ejemplares similares
-
Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction
por: Fan, Cunhang, et al.
Publicado: (2025) -
Low-latency auditory spatial attention detection based on spectro-spatial features from EEG
por: Cai, Siqi, et al.
Publicado: (2021) -
ExPO: Explainable Phonetic Trait-Oriented Network for Speaker Verification
por: Ma, Yi, et al.
Publicado: (2025) -
Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction
por: Pan, Zexu, et al.
Publicado: (2025) -
Neuro-MSBG: An End-to-End Neural Model for Hearing Loss Simulation
por: Yuan, Hui-Guan, et al.
Publicado: (2025)