Speech Representation Analysis based on Inter- and Intra-Model Similarities
Fuente:
arXiv
Saved in:
| Main Authors: | Kheir, Yassine El, Ali, Ahmed, Chowdhury, Shammur Absar |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IQRA 2026: Interspeech Challenge on Automatic Pronunciation Assessment for Modern Standard Arabic (MSA)
by: Kheir, Yassine El, et al.
Published: (2026)
by: Kheir, Yassine El, et al.
Published: (2026)
Beyond Orthography: Automatic Recovery of Short Vowels and Dialectal Sounds in Arabic
by: Kheir, Yassine El, et al.
Published: (2024)
by: Kheir, Yassine El, et al.
Published: (2024)
Children's Speech Recognition through Discrete Token Enhancement
by: Sukhadia, Vrunda N., et al.
Published: (2024)
by: Sukhadia, Vrunda N., et al.
Published: (2024)
CAFE A Novel Code switching Dataset for Algerian Dialect French and English
by: Lachemat, Houssam Eddine-Othman, et al.
Published: (2024)
by: Lachemat, Houssam Eddine-Othman, et al.
Published: (2024)
BiCrossMamba-ST: Speech Deepfake Detection with Bidirectional Mamba Spectro-Temporal Cross-Attention
by: Kheir, Yassine El, et al.
Published: (2025)
by: Kheir, Yassine El, et al.
Published: (2025)
Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection
by: Kheir, Yassine El, et al.
Published: (2025)
by: Kheir, Yassine El, et al.
Published: (2025)
Generalizable Audio Spoofing Detection using Non-Semantic Representations
by: Das, Arnab, et al.
Published: (2025)
by: Das, Arnab, et al.
Published: (2025)
Intra- and Inter-modal Context Interaction Modeling for Conversational Speech Synthesis
by: Jia, Zhenqi, et al.
Published: (2024)
by: Jia, Zhenqi, et al.
Published: (2024)
Two Views, One Truth: Spectral and Self-Supervised Features Fusion for Robust Speech Deepfake Detection
by: Kheir, Yassine El, et al.
Published: (2025)
by: Kheir, Yassine El, et al.
Published: (2025)
HARNESS: Lightweight Distilled Arabic Speech Foundation Models
by: Sukhadia, Vrunda N., et al.
Published: (2026)
by: Sukhadia, Vrunda N., et al.
Published: (2026)
Learning Separated Representations for Instrument-based Music Similarity
by: Hashizume, Yuka, et al.
Published: (2025)
by: Hashizume, Yuka, et al.
Published: (2025)
SECodec: Structural Entropy-based Compressive Speech Representation Codec for Speech Language Models
by: Wang, Linqin, et al.
Published: (2024)
by: Wang, Linqin, et al.
Published: (2024)
DeepFense: A Unified, Modular, and Extensible Framework for Robust Deepfake Audio Detection
by: Kheir, Yassine El, et al.
Published: (2026)
by: Kheir, Yassine El, et al.
Published: (2026)
Towards a Unified Benchmark for Arabic Pronunciation Assessment: Quranic Recitation as Case Study
by: Kheir, Yassine El, et al.
Published: (2025)
by: Kheir, Yassine El, et al.
Published: (2025)
From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models
by: Ersoy, Asım, et al.
Published: (2025)
by: Ersoy, Asım, et al.
Published: (2025)
Inter-Speaker Relative Cues for Text-Guided Target Speech Extraction
by: Dai, Wang, et al.
Published: (2025)
by: Dai, Wang, et al.
Published: (2025)
Geometric Analysis of Speech Representation Spaces: Topological Disentanglement and Confound Detection
by: Kashyap, Bipasha, et al.
Published: (2026)
by: Kashyap, Bipasha, et al.
Published: (2026)
Assessing the Alignment of Audio Representations with Timbre Similarity Ratings
by: Tian, Haokun, et al.
Published: (2025)
by: Tian, Haokun, et al.
Published: (2025)
Progressive Residual Extraction based Pre-training for Speech Representation Learning
by: Wang, Tianrui, et al.
Published: (2024)
by: Wang, Tianrui, et al.
Published: (2024)
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
by: Alsayegh, Ali, et al.
Published: (2025)
by: Alsayegh, Ali, et al.
Published: (2025)
SVSNet+: Enhancing Speaker Voice Similarity Assessment Models with Representations from Speech Foundation Models
by: Yin, Chun, et al.
Published: (2024)
by: Yin, Chun, et al.
Published: (2024)
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
by: Li, Jialu, et al.
Published: (2024)
by: Li, Jialu, et al.
Published: (2024)
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
by: Ji, Shengpeng, et al.
Published: (2024)
by: Ji, Shengpeng, et al.
Published: (2024)
Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning
by: Ma, Ding, et al.
Published: (2026)
by: Ma, Ding, et al.
Published: (2026)
Learning Multidimensional Disentangled Representations of Instrumental Sounds for Musical Similarity Assessment
by: Hashizume, Yuka, et al.
Published: (2024)
by: Hashizume, Yuka, et al.
Published: (2024)
Simulating Native Speaker Shadowing for Nonnative Speech Assessment with Latent Speech Representations
by: Geng, Haopeng, et al.
Published: (2024)
by: Geng, Haopeng, et al.
Published: (2024)
A Large-Scale Probing Analysis of Speaker-Specific Attributes in Self-Supervised Speech Representations
by: Chiu, Aemon Yat Fei, et al.
Published: (2025)
by: Chiu, Aemon Yat Fei, et al.
Published: (2025)
Mamba-based Decoder-Only Approach with Bidirectional Speech Modeling for Speech Recognition
by: Masuyama, Yoshiki, et al.
Published: (2024)
by: Masuyama, Yoshiki, et al.
Published: (2024)
Improvement Speaker Similarity for Zero-Shot Any-to-Any Voice Conversion of Whispered and Regular Speech
by: Avdeeva, Anastasia, et al.
Published: (2024)
by: Avdeeva, Anastasia, et al.
Published: (2024)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
by: Wang, Shih-heng, et al.
Published: (2024)
by: Wang, Shih-heng, et al.
Published: (2024)
Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
by: Deng, Yimin, et al.
Published: (2024)
by: Deng, Yimin, et al.
Published: (2024)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
by: Han, Seungu, et al.
Published: (2026)
by: Han, Seungu, et al.
Published: (2026)
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
by: Zhang, Jing-Xuan, et al.
Published: (2025)
by: Zhang, Jing-Xuan, et al.
Published: (2025)
Deep Speech Synthesis from Multimodal Articulatory Representations
by: Wu, Peter, et al.
Published: (2024)
by: Wu, Peter, et al.
Published: (2024)
Convexity-based Pruning of Speech Representation Models
by: Dorszewski, Teresa, et al.
Published: (2024)
by: Dorszewski, Teresa, et al.
Published: (2024)
Music Similarity Representation Learning Focusing on Individual Instruments with Source Separation and Human Preference
by: Imamura, Takehiro, et al.
Published: (2025)
by: Imamura, Takehiro, et al.
Published: (2025)
Adaptive Convolution for CNN-based Speech Enhancement Models
by: Wang, Dahan, et al.
Published: (2025)
by: Wang, Dahan, et al.
Published: (2025)
SingOMD: Singing Oriented Multi-resolution Discrete Representation Construction from Speech Models
by: Tang, Yuxun, et al.
Published: (2024)
by: Tang, Yuxun, et al.
Published: (2024)
Analysing the Masked predictive coding training criterion for pre-training a Speech Representation Model
by: Yadav, Hemant, et al.
Published: (2023)
by: Yadav, Hemant, et al.
Published: (2023)
CVSM: Contrastive Vocal Similarity Modeling
by: Garoufis, Christos, et al.
Published: (2025)
by: Garoufis, Christos, et al.
Published: (2025)
Similar Items
-
IQRA 2026: Interspeech Challenge on Automatic Pronunciation Assessment for Modern Standard Arabic (MSA)
by: Kheir, Yassine El, et al.
Published: (2026) -
Beyond Orthography: Automatic Recovery of Short Vowels and Dialectal Sounds in Arabic
by: Kheir, Yassine El, et al.
Published: (2024) -
Children's Speech Recognition through Discrete Token Enhancement
by: Sukhadia, Vrunda N., et al.
Published: (2024) -
CAFE A Novel Code switching Dataset for Algerian Dialect French and English
by: Lachemat, Houssam Eddine-Othman, et al.
Published: (2024) -
BiCrossMamba-ST: Speech Deepfake Detection with Bidirectional Mamba Spectro-Temporal Cross-Attention
by: Kheir, Yassine El, et al.
Published: (2025)