What do Speech Foundation Models Learn? Analysis and Applications
Fuente:
arXiv
Salvato in:
| Autore principale: | Pasad, Ankita |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
What Do Self-Supervised Speech Models Know About Words?
di: Pasad, Ankita, et al.
Pubblicazione: (2023)
di: Pasad, Ankita, et al.
Pubblicazione: (2023)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
di: Arora, Siddhant, et al.
Pubblicazione: (2024)
di: Arora, Siddhant, et al.
Pubblicazione: (2024)
Training and Inference Efficiency of Encoder-Decoder Speech Models
di: Żelasko, Piotr, et al.
Pubblicazione: (2025)
di: Żelasko, Piotr, et al.
Pubblicazione: (2025)
Self-Supervised Speech Representations are More Phonetic than Semantic
di: Choi, Kwanghee, et al.
Pubblicazione: (2024)
di: Choi, Kwanghee, et al.
Pubblicazione: (2024)
What Do Speech Foundation Models Not Learn About Speech?
di: Waheed, Abdul, et al.
Pubblicazione: (2024)
di: Waheed, Abdul, et al.
Pubblicazione: (2024)
POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
di: Li, Chin-Jou, et al.
Pubblicazione: (2025)
di: Li, Chin-Jou, et al.
Pubblicazione: (2025)
Hallucination Benchmark for Speech Foundation Models
di: Koudounas, Alkis, et al.
Pubblicazione: (2025)
di: Koudounas, Alkis, et al.
Pubblicazione: (2025)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
di: Shivakumar, Prashanth Gurunath, et al.
Pubblicazione: (2024)
di: Shivakumar, Prashanth Gurunath, et al.
Pubblicazione: (2024)
A Large-Scale Evaluation of Speech Foundation Models
di: Yang, Shu-wen, et al.
Pubblicazione: (2024)
di: Yang, Shu-wen, et al.
Pubblicazione: (2024)
What Do Self-Supervised Speech and Speaker Models Learn? New Findings From a Cross Model Layer-Wise Analysis
di: Ashihara, Takanori, et al.
Pubblicazione: (2024)
di: Ashihara, Takanori, et al.
Pubblicazione: (2024)
On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
di: Tian, Jinchuan, et al.
Pubblicazione: (2024)
di: Tian, Jinchuan, et al.
Pubblicazione: (2024)
LESS: Large Language Model Enhanced Semi-Supervised Learning for Speech Foundational Models Using in-the-wild Data
di: Ding, Wen, et al.
Pubblicazione: (2025)
di: Ding, Wen, et al.
Pubblicazione: (2025)
Transducer Consistency Regularization for Speech to Text Applications
di: Tseng, Cindy, et al.
Pubblicazione: (2024)
di: Tseng, Cindy, et al.
Pubblicazione: (2024)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
di: Peng, Yifan, et al.
Pubblicazione: (2024)
di: Peng, Yifan, et al.
Pubblicazione: (2024)
Speech Foundation Models and Crowdsourcing for Efficient, High-Quality Data Collection
di: Lee, Beomseok, et al.
Pubblicazione: (2024)
di: Lee, Beomseok, et al.
Pubblicazione: (2024)
Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
di: Fan, Ruchao, et al.
Pubblicazione: (2024)
di: Fan, Ruchao, et al.
Pubblicazione: (2024)
iMiGUE-Speech: A Spontaneous Speech Dataset for Affective Analysis
di: Kakouros, Sofoklis, et al.
Pubblicazione: (2026)
di: Kakouros, Sofoklis, et al.
Pubblicazione: (2026)
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
di: Raina, Vyas, et al.
Pubblicazione: (2024)
di: Raina, Vyas, et al.
Pubblicazione: (2024)
Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
di: Raina, Vyas, et al.
Pubblicazione: (2024)
di: Raina, Vyas, et al.
Pubblicazione: (2024)
Unifying Model and Layer Fusion for Speech Foundation Models
di: Shih, Yi-Jen, et al.
Pubblicazione: (2025)
di: Shih, Yi-Jen, et al.
Pubblicazione: (2025)
Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
di: Feng, Tiantian, et al.
Pubblicazione: (2025)
di: Feng, Tiantian, et al.
Pubblicazione: (2025)
Investigating Zero-Shot Generalizability on Mandarin-English Code-Switched ASR and Speech-to-text Translation of Recent Foundation Models with Self-Supervision and Weak Supervision
di: Yang, Chih-Kai, et al.
Pubblicazione: (2023)
di: Yang, Chih-Kai, et al.
Pubblicazione: (2023)
SpeechCT-CLIP: Distilling Text-Image Knowledge to Speech for Voice-Native Multimodal CT Analysis
di: Buess, Lukas, et al.
Pubblicazione: (2025)
di: Buess, Lukas, et al.
Pubblicazione: (2025)
AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation
di: Manakul, Potsawee, et al.
Pubblicazione: (2025)
di: Manakul, Potsawee, et al.
Pubblicazione: (2025)
HARNESS: Lightweight Distilled Arabic Speech Foundation Models
di: Sukhadia, Vrunda N., et al.
Pubblicazione: (2026)
di: Sukhadia, Vrunda N., et al.
Pubblicazione: (2026)
What Are They Doing? Joint Audio-Speech Co-Reasoning
di: Wang, Yingzhi, et al.
Pubblicazione: (2024)
di: Wang, Yingzhi, et al.
Pubblicazione: (2024)
Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2025)
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2025)
MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond
di: Huzaifah, Muhammad, et al.
Pubblicazione: (2024)
di: Huzaifah, Muhammad, et al.
Pubblicazione: (2024)
Linguistic Knowledge Transfer Learning for Speech Enhancement
di: Hung, Kuo-Hsuan, et al.
Pubblicazione: (2025)
di: Hung, Kuo-Hsuan, et al.
Pubblicazione: (2025)
Learning Speech Representations with Variational Predictive Coding
di: Yeh, Sung-Lin, et al.
Pubblicazione: (2025)
di: Yeh, Sung-Lin, et al.
Pubblicazione: (2025)
Towards Efficient Speech-Text Jointly Decoding within One Speech Language Model
di: Wu, Haibin, et al.
Pubblicazione: (2025)
di: Wu, Haibin, et al.
Pubblicazione: (2025)
Speech Recognition Model Improves Text-to-Speech Synthesis using Fine-Grained Reward
di: Wang, Guansu, et al.
Pubblicazione: (2025)
di: Wang, Guansu, et al.
Pubblicazione: (2025)
DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment
di: Lu, Ke-Han, et al.
Pubblicazione: (2024)
di: Lu, Ke-Han, et al.
Pubblicazione: (2024)
Automatic Speech Recognition of Non-Native Child Speech for Language Learning Applications
di: Wills, Simone, et al.
Pubblicazione: (2023)
di: Wills, Simone, et al.
Pubblicazione: (2023)
A Preliminary Analysis of Automatic Word and Syllable Prominence Detection in Non-Native Speech With Text-to-Speech Prosody Embeddings
di: Mondal, Anindita, et al.
Pubblicazione: (2024)
di: Mondal, Anindita, et al.
Pubblicazione: (2024)
Scaling Analysis of Interleaved Speech-Text Language Models
di: Maimon, Gallil, et al.
Pubblicazione: (2025)
di: Maimon, Gallil, et al.
Pubblicazione: (2025)
Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models
di: Lu, Ke-Han, et al.
Pubblicazione: (2025)
di: Lu, Ke-Han, et al.
Pubblicazione: (2025)
MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model
di: Park, Joonyong, et al.
Pubblicazione: (2025)
di: Park, Joonyong, et al.
Pubblicazione: (2025)
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition
di: Tseng, Yuan, et al.
Pubblicazione: (2025)
di: Tseng, Yuan, et al.
Pubblicazione: (2025)
SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style Captioning
di: Huang, Chien-yu, et al.
Pubblicazione: (2024)
di: Huang, Chien-yu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
What Do Self-Supervised Speech Models Know About Words?
di: Pasad, Ankita, et al.
Pubblicazione: (2023) -
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
di: Arora, Siddhant, et al.
Pubblicazione: (2024) -
Training and Inference Efficiency of Encoder-Decoder Speech Models
di: Żelasko, Piotr, et al.
Pubblicazione: (2025) -
Self-Supervised Speech Representations are More Phonetic than Semantic
di: Choi, Kwanghee, et al.
Pubblicazione: (2024) -
What Do Speech Foundation Models Not Learn About Speech?
di: Waheed, Abdul, et al.
Pubblicazione: (2024)