What Do Self-Supervised Speech Models Know About Words?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pasad, Ankita, Chien, Chung-Ming, Settle, Shane, Livescu, Karen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Supervised Speech Representations are More Phonetic than Semantic
von: Choi, Kwanghee, et al.
Veröffentlicht: (2024)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2024)
What do Speech Foundation Models Learn? Analysis and Applications
von: Pasad, Ankita
Veröffentlicht: (2025)
von: Pasad, Ankita
Veröffentlicht: (2025)
AV2Wav: Diffusion-Based Re-synthesis from Continuous Self-supervised Features for Audio-Visual Speech Enhancement
von: Chou, Ju-Chieh, et al.
Veröffentlicht: (2023)
von: Chou, Ju-Chieh, et al.
Veröffentlicht: (2023)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models
von: Chien, Chung-Ming, et al.
Veröffentlicht: (2026)
von: Chien, Chung-Ming, et al.
Veröffentlicht: (2026)
CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition
von: Bartelds, Martijn, et al.
Veröffentlicht: (2025)
von: Bartelds, Martijn, et al.
Veröffentlicht: (2025)
Integrating Self-supervised Speech Model with Pseudo Word-level Targets from Visually-grounded Speech Model
von: Fang, Hung-Chieh, et al.
Veröffentlicht: (2024)
von: Fang, Hung-Chieh, et al.
Veröffentlicht: (2024)
DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding
von: Shon, Suwon, et al.
Veröffentlicht: (2024)
von: Shon, Suwon, et al.
Veröffentlicht: (2024)
What Do Speech Foundation Models Not Learn About Speech?
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
Towards Early Prediction of Self-Supervised Speech Model Performance
von: Whetten, Ryan, et al.
Veröffentlicht: (2025)
von: Whetten, Ryan, et al.
Veröffentlicht: (2025)
Is Smaller Always Faster? Tradeoffs in Compressing Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
Property Neurons in Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2024)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2024)
Enhancing the Stability of LLM-based Speech Generation Systems through Self-Supervised Representations
von: Martín-Cortinas, Álvaro, et al.
Veröffentlicht: (2024)
von: Martín-Cortinas, Álvaro, et al.
Veröffentlicht: (2024)
Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
von: Liu, Andy T., et al.
Veröffentlicht: (2024)
von: Liu, Andy T., et al.
Veröffentlicht: (2024)
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
Training and Inference Efficiency of Encoder-Decoder Speech Models
von: Żelasko, Piotr, et al.
Veröffentlicht: (2025)
von: Żelasko, Piotr, et al.
Veröffentlicht: (2025)
Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
von: Choi, Kwanghee, et al.
Veröffentlicht: (2025)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2025)
Task Oriented Dialogue as a Catalyst for Self-Supervised Automatic Speech Recognition
von: Chan, David M., et al.
Veröffentlicht: (2024)
von: Chan, David M., et al.
Veröffentlicht: (2024)
SKILL: Similarity-aware Knowledge distILLation for Speech Self-Supervised Learning
von: Zampierin, Luca, et al.
Veröffentlicht: (2024)
von: Zampierin, Luca, et al.
Veröffentlicht: (2024)
Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing
von: Trinh, Viet Anh, et al.
Veröffentlicht: (2024)
von: Trinh, Viet Anh, et al.
Veröffentlicht: (2024)
CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2024)
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2024)
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
von: Fan, Xulin, et al.
Veröffentlicht: (2026)
von: Fan, Xulin, et al.
Veröffentlicht: (2026)
Universal Robust Speech Adaptation for Cross-Domain Speech Recognition and Enhancement
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2026)
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2026)
Africa-Centric Self-Supervised Pre-Training for Multilingual Speech Representation in a Sub-Saharan Context
von: Caubrière, Antoine, et al.
Veröffentlicht: (2024)
von: Caubrière, Antoine, et al.
Veröffentlicht: (2024)
On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
von: Tian, Jinchuan, et al.
Veröffentlicht: (2024)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2024)
Language Bias in Self-Supervised Learning For Automatic Speech Recognition
von: Storey, Edward, et al.
Veröffentlicht: (2025)
von: Storey, Edward, et al.
Veröffentlicht: (2025)
Comparing Self-Supervised Learning Models Pre-Trained on Human Speech and Animal Vocalizations for Bioacoustics Processing
von: Sarkar, Eklavya, et al.
Veröffentlicht: (2025)
von: Sarkar, Eklavya, et al.
Veröffentlicht: (2025)
TaigiSpeech: A Low-Resource Real-World Speech Intent Dataset and Preliminary Results with Scalable Data Mining In-the-Wild
von: Chang, Kai-Wei, et al.
Veröffentlicht: (2026)
von: Chang, Kai-Wei, et al.
Veröffentlicht: (2026)
Data-Centric Lessons To Improve Speech-Language Pretraining
von: Udandarao, Vishaal, et al.
Veröffentlicht: (2025)
von: Udandarao, Vishaal, et al.
Veröffentlicht: (2025)
DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
von: Chang, Heng-Jui, et al.
Veröffentlicht: (2024)
von: Chang, Heng-Jui, et al.
Veröffentlicht: (2024)
RosettaSpeech: Zero-Shot Speech-to-Speech Translation without Parallel Speech
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
What Do Self-Supervised Speech and Speaker Models Learn? New Findings From a Cross Model Layer-Wise Analysis
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024)
Gender Bias in Instruction-Guided Speech Synthesis Models
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2025)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2025)
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
Improving Speech Decoding from ECoG with Self-Supervised Pretraining
von: Yuan, Brian A., et al.
Veröffentlicht: (2024)
von: Yuan, Brian A., et al.
Veröffentlicht: (2024)
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
von: Park, Chanho, et al.
Veröffentlicht: (2023)
von: Park, Chanho, et al.
Veröffentlicht: (2023)
Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
von: Park, Taejin, et al.
Veröffentlicht: (2024)
von: Park, Taejin, et al.
Veröffentlicht: (2024)
A Contrastive Learning Approach to Mitigate Bias in Speech Models
von: Koudounas, Alkis, et al.
Veröffentlicht: (2024)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Self-Supervised Speech Representations are More Phonetic than Semantic
von: Choi, Kwanghee, et al.
Veröffentlicht: (2024) -
What do Speech Foundation Models Learn? Analysis and Applications
von: Pasad, Ankita
Veröffentlicht: (2025) -
AV2Wav: Diffusion-Based Re-synthesis from Continuous Self-supervised Features for Audio-Visual Speech Enhancement
von: Chou, Ju-Chieh, et al.
Veröffentlicht: (2023) -
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
von: Arora, Siddhant, et al.
Veröffentlicht: (2024) -
MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models
von: Chien, Chung-Ming, et al.
Veröffentlicht: (2026)