Self-supervised Speech Models for Word-Level Stuttered Speech Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shih, Yi-Jen, Gkalitsiou, Zoi, Dimakis, Alexandros G., Harwath, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Interface Design for Self-Supervised Speech Models
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
Unifying Model and Layer Fusion for Speech Foundation Models
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2025)
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2025)
ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining
von: Diwan, Anuj, et al.
Veröffentlicht: (2026)
von: Diwan, Anuj, et al.
Veröffentlicht: (2026)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
Large Language Models for Dysfluency Detection in Stuttered Speech
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
Scaling Rich Style-Prompted Text-to-Speech Datasets
von: Diwan, Anuj, et al.
Veröffentlicht: (2025)
von: Diwan, Anuj, et al.
Veröffentlicht: (2025)
VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
von: Peng, Puyuan, et al.
Veröffentlicht: (2024)
von: Peng, Puyuan, et al.
Veröffentlicht: (2024)
Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
Multilingual Stutter Event Detection for English, German, and Mandarin Speech
von: Haas, Felix, et al.
Veröffentlicht: (2026)
von: Haas, Felix, et al.
Veröffentlicht: (2026)
Deploying UDM Series in Real-Life Stuttered Speech Applications: A Clinical Evaluation Framework
von: Zhang, Eric, et al.
Veröffentlicht: (2025)
von: Zhang, Eric, et al.
Veröffentlicht: (2025)
From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models
von: Ersoy, Asım, et al.
Veröffentlicht: (2025)
von: Ersoy, Asım, et al.
Veröffentlicht: (2025)
LeBenchmark 2.0: a Standardized, Replicable and Enhanced Framework for Self-supervised Representations of French Speech
von: Parcollet, Titouan, et al.
Veröffentlicht: (2023)
von: Parcollet, Titouan, et al.
Veröffentlicht: (2023)
Harf-Speech: A Clinically Aligned Framework for Arabic Phoneme-Level Speech Assessment
von: Azad, Asif, et al.
Veröffentlicht: (2026)
von: Azad, Asif, et al.
Veröffentlicht: (2026)
Hierarchical Self-Supervised Representation Learning for Depression Detection from Speech
von: Li, Yuxin, et al.
Veröffentlicht: (2025)
von: Li, Yuxin, et al.
Veröffentlicht: (2025)
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
SeamlessExpressiveLM: Speech Language Model for Expressive Speech-to-Speech Translation with Chain-of-Thought
von: Gong, Hongyu, et al.
Veröffentlicht: (2024)
von: Gong, Hongyu, et al.
Veröffentlicht: (2024)
Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations
von: Meghanani, Amit, et al.
Veröffentlicht: (2024)
von: Meghanani, Amit, et al.
Veröffentlicht: (2024)
Probing the Robustness Properties of Neural Speech Codecs
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
Device-Directed Speech Detection for Follow-up Conversations Using Large Language Models
von: Ognjen, et al.
Veröffentlicht: (2024)
von: Ognjen, et al.
Veröffentlicht: (2024)
StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024)
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024)
SSDM: Scalable Speech Dysfluency Modeling
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
SPBA: Utilizing Speech Large Language Model for Backdoor Attacks on Speech Classification Models
von: Yao, Wenhan, et al.
Veröffentlicht: (2025)
von: Yao, Wenhan, et al.
Veröffentlicht: (2025)
Adapting Foundation Speech Recognition Models to Impaired Speech: A Semantic Re-chaining Approach for Personalization of German Speech
von: Pokel, Niclas, et al.
Veröffentlicht: (2025)
von: Pokel, Niclas, et al.
Veröffentlicht: (2025)
A Benchmark for Early-stage Parkinson's Disease Detection from Speech
von: Zhong, Terry Yi, et al.
Veröffentlicht: (2026)
von: Zhong, Terry Yi, et al.
Veröffentlicht: (2026)
An Investigation Into Explainable Audio Hate Speech Detection
von: An, Jinmyeong, et al.
Veröffentlicht: (2024)
von: An, Jinmyeong, et al.
Veröffentlicht: (2024)
Quantizer-Aware Hierarchical Neural Codec Modeling for Speech Deepfake Detection
von: Wu, Jinyang, et al.
Veröffentlicht: (2026)
von: Wu, Jinyang, et al.
Veröffentlicht: (2026)
Integrating Self-supervised Speech Model with Pseudo Word-level Targets from Visually-grounded Speech Model
von: Fang, Hung-Chieh, et al.
Veröffentlicht: (2024)
von: Fang, Hung-Chieh, et al.
Veröffentlicht: (2024)
Stutter-Solver: End-to-end Multi-lingual Dysfluency Detection
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations
von: Meghanani, Amit, et al.
Veröffentlicht: (2026)
von: Meghanani, Amit, et al.
Veröffentlicht: (2026)
MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation
von: Chen, Szu-Chi, et al.
Veröffentlicht: (2026)
von: Chen, Szu-Chi, et al.
Veröffentlicht: (2026)
Dynamic Stress Detection: A Study of Temporal Progression Modelling of Stress in Speech
von: Lall, Vishakha, et al.
Veröffentlicht: (2025)
von: Lall, Vishakha, et al.
Veröffentlicht: (2025)
Towards Improving NAM-to-Speech Synthesis Intelligibility using Self-Supervised Speech Models
von: Shah, Neil, et al.
Veröffentlicht: (2024)
von: Shah, Neil, et al.
Veröffentlicht: (2024)
Temporal-Channel Modeling in Multi-head Self-Attention for Synthetic Speech Detection
von: Truong, Duc-Tuan, et al.
Veröffentlicht: (2024)
von: Truong, Duc-Tuan, et al.
Veröffentlicht: (2024)
BAT: Learning to Reason about Spatial Sounds with Large Language Models
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2024)
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2024)
Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis
von: Ye, Zongli, et al.
Veröffentlicht: (2025)
von: Ye, Zongli, et al.
Veröffentlicht: (2025)
SpeechComposer: Unifying Multiple Speech Tasks with Prompt Composition
von: Wu, Yihan, et al.
Veröffentlicht: (2024)
von: Wu, Yihan, et al.
Veröffentlicht: (2024)
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
Neural networks for Text-to-Speech evaluation
von: Trofimenko, Ilya, et al.
Veröffentlicht: (2026)
von: Trofimenko, Ilya, et al.
Veröffentlicht: (2026)
SpeechPrune: Context-aware Token Pruning for Speech Information Retrieval
von: Lin, Yueqian, et al.
Veröffentlicht: (2024)
von: Lin, Yueqian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Interface Design for Self-Supervised Speech Models
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024) -
Unifying Model and Layer Fusion for Speech Foundation Models
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2025) -
ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining
von: Diwan, Anuj, et al.
Veröffentlicht: (2026) -
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024) -
Large Language Models for Dysfluency Detection in Stuttered Speech
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)