Geometric Analysis of Speech Representation Spaces: Topological Disentanglement and Confound Detection
Fuente:
arXiv
Guardado en:
| Autores principales: | Kashyap, Bipasha, Pathirana, Pubudu N. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Quantifying Dimensional Independence in Speech: An Information-Theoretic Framework for Disentangled Representation Learning
por: Kashyap, Bipasha, et al.
Publicado: (2026)
por: Kashyap, Bipasha, et al.
Publicado: (2026)
Quantifying Quanvolutional Neural Networks Robustness for Speech in Healthcare Applications
por: Tran, Ha, et al.
Publicado: (2026)
por: Tran, Ha, et al.
Publicado: (2026)
Quantum-Inspired Genetic Algorithm for Robust Source Separation in Smart City Acoustics
por: Quan, Minh K., et al.
Publicado: (2025)
por: Quan, Minh K., et al.
Publicado: (2025)
Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
por: Deng, Yimin, et al.
Publicado: (2024)
por: Deng, Yimin, et al.
Publicado: (2024)
Learning Disentangled Speech Representations with Contrastive Learning and Time-Invariant Retrieval
por: Deng, Yimin, et al.
Publicado: (2024)
por: Deng, Yimin, et al.
Publicado: (2024)
Learning Disentangled Speech Representations
por: Brima, Yusuf, et al.
Publicado: (2023)
por: Brima, Yusuf, et al.
Publicado: (2023)
Quantum-Enhanced Transformers for Robust Acoustic Scene Classification in IoT Environments
por: Quan, Minh K., et al.
Publicado: (2025)
por: Quan, Minh K., et al.
Publicado: (2025)
Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance
por: Ochiai, Tsubasa, et al.
Publicado: (2024)
por: Ochiai, Tsubasa, et al.
Publicado: (2024)
HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement
por: Hussein, Amir, et al.
Publicado: (2025)
por: Hussein, Amir, et al.
Publicado: (2025)
Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability
por: Zhu, Xiaoxu, et al.
Publicado: (2025)
por: Zhu, Xiaoxu, et al.
Publicado: (2025)
Unified Architecture and Unsupervised Speech Disentanglement for Speaker Embedding-Free Enrollment in Personalized Speech Enhancement
por: Huang, Ziling, et al.
Publicado: (2025)
por: Huang, Ziling, et al.
Publicado: (2025)
Speech Representation Analysis based on Inter- and Intra-Model Similarities
por: Kheir, Yassine El, et al.
Publicado: (2024)
por: Kheir, Yassine El, et al.
Publicado: (2024)
Disentangled Representation Learning for Environment-agnostic Speaker Recognition
por: Nam, KiHyun, et al.
Publicado: (2024)
por: Nam, KiHyun, et al.
Publicado: (2024)
Noisy Disentanglement with Tri-stage Training for Noise-Robust Speech Recognition
por: Chen, Shuangyuan, et al.
Publicado: (2025)
por: Chen, Shuangyuan, et al.
Publicado: (2025)
Are Paralinguistic Representations all that is needed for Speech Emotion Recognition?
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
por: Melechovsky, Jan, et al.
Publicado: (2024)
por: Melechovsky, Jan, et al.
Publicado: (2024)
Comparative Analysis of ASR Methods for Speech Deepfake Detection
por: Salvi, Davide, et al.
Publicado: (2024)
por: Salvi, Davide, et al.
Publicado: (2024)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
por: Lu, Ye-Xin, et al.
Publicado: (2025)
por: Lu, Ye-Xin, et al.
Publicado: (2025)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
por: Sato, Hiroshi, et al.
Publicado: (2025)
por: Sato, Hiroshi, et al.
Publicado: (2025)
Learning Multidimensional Disentangled Representations of Instrumental Sounds for Musical Similarity Assessment
por: Hashizume, Yuka, et al.
Publicado: (2024)
por: Hashizume, Yuka, et al.
Publicado: (2024)
Balancing Information Preservation and Disentanglement in Self-Supervised Music Representation Learning
por: Wilkins, Julia, et al.
Publicado: (2025)
por: Wilkins, Julia, et al.
Publicado: (2025)
Self-Supervised Multi-View Learning for Disentangled Music Audio Representations
por: Wilkins, Julia, et al.
Publicado: (2024)
por: Wilkins, Julia, et al.
Publicado: (2024)
Phoneme-Level Analysis for Person-of-Interest Speech Deepfake Detection
por: Salvi, Davide, et al.
Publicado: (2025)
por: Salvi, Davide, et al.
Publicado: (2025)
A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
por: Pham, Lam, et al.
Publicado: (2024)
por: Pham, Lam, et al.
Publicado: (2024)
Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning
por: Ma, Ding, et al.
Publicado: (2026)
por: Ma, Ding, et al.
Publicado: (2026)
FRCRN: Boosting Feature Representation using Frequency Recurrence for Monaural Speech Enhancement
por: Zhao, Shengkui, et al.
Publicado: (2022)
por: Zhao, Shengkui, et al.
Publicado: (2022)
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
por: Zhou, Xuehao, et al.
Publicado: (2024)
por: Zhou, Xuehao, et al.
Publicado: (2024)
DIVINE: Coordinating Multimodal Disentangled Representations for Oro-Facial Neurological Disorder Assessment
por: Akhtar, Mohd Mujtaba, et al.
Publicado: (2026)
por: Akhtar, Mohd Mujtaba, et al.
Publicado: (2026)
Simulating Native Speaker Shadowing for Nonnative Speech Assessment with Latent Speech Representations
por: Geng, Haopeng, et al.
Publicado: (2024)
por: Geng, Haopeng, et al.
Publicado: (2024)
A Large-Scale Probing Analysis of Speaker-Specific Attributes in Self-Supervised Speech Representations
por: Chiu, Aemon Yat Fei, et al.
Publicado: (2025)
por: Chiu, Aemon Yat Fei, et al.
Publicado: (2025)
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
por: Gao, Xiaoxue, et al.
Publicado: (2024)
por: Gao, Xiaoxue, et al.
Publicado: (2024)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
por: Wang, Shih-heng, et al.
Publicado: (2024)
por: Wang, Shih-heng, et al.
Publicado: (2024)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
por: Han, Seungu, et al.
Publicado: (2026)
por: Han, Seungu, et al.
Publicado: (2026)
SECodec: Structural Entropy-based Compressive Speech Representation Codec for Speech Language Models
por: Wang, Linqin, et al.
Publicado: (2024)
por: Wang, Linqin, et al.
Publicado: (2024)
Audio-text Retrieval with Transformer-based Hierarchical Alignment and Disentangled Cross-modal Representation
por: Xin, Yifei, et al.
Publicado: (2024)
por: Xin, Yifei, et al.
Publicado: (2024)
Evaluating Parkinson's Disease Detection in Anonymized Speech: A Performance and Acoustic Analysis
por: Franzreb, Carlos, et al.
Publicado: (2026)
por: Franzreb, Carlos, et al.
Publicado: (2026)
Deep Speech Synthesis from Multimodal Articulatory Representations
por: Wu, Peter, et al.
Publicado: (2024)
por: Wu, Peter, et al.
Publicado: (2024)
Disentangling Hierarchical Features for Anomalous Sound Detection Under Domain Shift
por: Guan, Jian, et al.
Publicado: (2025)
por: Guan, Jian, et al.
Publicado: (2025)
ParaMETA: Towards Learning Disentangled Paralinguistic Speaking Styles Representations from Speech
por: Lou, Haowei, et al.
Publicado: (2026)
por: Lou, Haowei, et al.
Publicado: (2026)
EAD-VC: Enhancing Speech Auto-Disentanglement for Voice Conversion with IFUB Estimator and Joint Text-Guided Consistent Learning
por: Liang, Ziqi, et al.
Publicado: (2024)
por: Liang, Ziqi, et al.
Publicado: (2024)
Ejemplares similares
-
Quantifying Dimensional Independence in Speech: An Information-Theoretic Framework for Disentangled Representation Learning
por: Kashyap, Bipasha, et al.
Publicado: (2026) -
Quantifying Quanvolutional Neural Networks Robustness for Speech in Healthcare Applications
por: Tran, Ha, et al.
Publicado: (2026) -
Quantum-Inspired Genetic Algorithm for Robust Source Separation in Smart City Acoustics
por: Quan, Minh K., et al.
Publicado: (2025) -
Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
por: Deng, Yimin, et al.
Publicado: (2024) -
Learning Disentangled Speech Representations with Contrastive Learning and Time-Invariant Retrieval
por: Deng, Yimin, et al.
Publicado: (2024)