Graph Modelling Analysis of Speech-Gesture Interaction for Aphasia Severity Estimation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kollapally, Navya Martin, Akers, Christa, Joseph, Renjith Nelson |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Voice Biomarker Analysis and Automated Severity Classification of Dysarthric Speech in a Multilingual Context
von: Yeo, Eunjung
Veröffentlicht: (2024)
von: Yeo, Eunjung
Veröffentlicht: (2024)
Addressing Pitfalls in Auditing Practices of Automatic Speech Recognition Technologies: A Case Study of People with Aphasia
von: Mei, Katelyn Xiaoying, et al.
Veröffentlicht: (2025)
von: Mei, Katelyn Xiaoying, et al.
Veröffentlicht: (2025)
Intra- and Inter-modal Context Interaction Modeling for Conversational Speech Synthesis
von: Jia, Zhenqi, et al.
Veröffentlicht: (2024)
von: Jia, Zhenqi, et al.
Veröffentlicht: (2024)
DiffAR: Denoising Diffusion Autoregressive Model for Raw Speech Waveform Generation
von: Benita, Roi, et al.
Veröffentlicht: (2023)
von: Benita, Roi, et al.
Veröffentlicht: (2023)
Scaling Analysis of Interleaved Speech-Text Language Models
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
von: Hu, Ke, et al.
Veröffentlicht: (2025)
von: Hu, Ke, et al.
Veröffentlicht: (2025)
Automatic Speech Recognition System-Independent Word Error Rate Estimation
von: Park, Chanho, et al.
Veröffentlicht: (2024)
von: Park, Chanho, et al.
Veröffentlicht: (2024)
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
What Do Speech Foundation Models Not Learn About Speech?
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
Unified Pathological Speech Analysis with Prompt Tuning
von: Yang, Fei, et al.
Veröffentlicht: (2024)
von: Yang, Fei, et al.
Veröffentlicht: (2024)
PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models
von: Yang, Runyan, et al.
Veröffentlicht: (2024)
von: Yang, Runyan, et al.
Veröffentlicht: (2024)
PredGen: Accelerated Inference of Large Language Models through Input-Time Speculation for Real-Time Speech Interaction
von: Li, Shufan, et al.
Veröffentlicht: (2025)
von: Li, Shufan, et al.
Veröffentlicht: (2025)
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
Compact Speech Translation Models via Discrete Speech Units Pretraining
von: Lam, Tsz Kin, et al.
Veröffentlicht: (2024)
von: Lam, Tsz Kin, et al.
Veröffentlicht: (2024)
An Empirical Study of Speech Language Models for Prompt-Conditioned Speech Synthesis
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models
von: Dowerah, Sandipana, et al.
Veröffentlicht: (2025)
von: Dowerah, Sandipana, et al.
Veröffentlicht: (2025)
Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
von: Kando, Shunsuke, et al.
Veröffentlicht: (2025)
von: Kando, Shunsuke, et al.
Veröffentlicht: (2025)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
von: Li, Zhu, et al.
Veröffentlicht: (2025)
von: Li, Zhu, et al.
Veröffentlicht: (2025)
MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
von: Park, Chanho, et al.
Veröffentlicht: (2023)
von: Park, Chanho, et al.
Veröffentlicht: (2023)
Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis
von: Xu, Tianyi, et al.
Veröffentlicht: (2025)
von: Xu, Tianyi, et al.
Veröffentlicht: (2025)
Psychoacoustic Challenges Of Speech Enhancement On VoIP Platforms
von: Konan, Joseph, et al.
Veröffentlicht: (2023)
von: Konan, Joseph, et al.
Veröffentlicht: (2023)
Effective Context in Neural Speech Models
von: Meng, Yen, et al.
Veröffentlicht: (2025)
von: Meng, Yen, et al.
Veröffentlicht: (2025)
Hallucination Benchmark for Speech Foundation Models
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations
von: Meghanani, Amit, et al.
Veröffentlicht: (2026)
von: Meghanani, Amit, et al.
Veröffentlicht: (2026)
Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities
von: Adila, Aulia, et al.
Veröffentlicht: (2024)
von: Adila, Aulia, et al.
Veröffentlicht: (2024)
EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
von: de Seyssel, Maureen, et al.
Veröffentlicht: (2023)
von: de Seyssel, Maureen, et al.
Veröffentlicht: (2023)
Age-Dependent Analysis and Stochastic Generation of Child-Directed Speech
von: Räsänen, Okko, et al.
Veröffentlicht: (2024)
von: Räsänen, Okko, et al.
Veröffentlicht: (2024)
TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
von: Peng, Junyi, et al.
Veröffentlicht: (2025)
von: Peng, Junyi, et al.
Veröffentlicht: (2025)
STaR: Distilling Speech Temporal Relation for Lightweight Speech Self-Supervised Learning Models
von: Jang, Kangwook, et al.
Veröffentlicht: (2023)
von: Jang, Kangwook, et al.
Veröffentlicht: (2023)
SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?
von: Ashihara, Takanori, et al.
Veröffentlicht: (2023)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2023)
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models
von: Fang, Yuanbo, et al.
Veröffentlicht: (2025)
von: Fang, Yuanbo, et al.
Veröffentlicht: (2025)
StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model
von: Guo, Shoutao, et al.
Veröffentlicht: (2025)
von: Guo, Shoutao, et al.
Veröffentlicht: (2025)
SEAL: Speech Embedding Alignment Learning for Speech Large Language Model with Retrieval-Augmented Generation
von: Sun, Chunyu, et al.
Veröffentlicht: (2025)
von: Sun, Chunyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Voice Biomarker Analysis and Automated Severity Classification of Dysarthric Speech in a Multilingual Context
von: Yeo, Eunjung
Veröffentlicht: (2024) -
Addressing Pitfalls in Auditing Practices of Automatic Speech Recognition Technologies: A Case Study of People with Aphasia
von: Mei, Katelyn Xiaoying, et al.
Veröffentlicht: (2025) -
Intra- and Inter-modal Context Interaction Modeling for Conversational Speech Synthesis
von: Jia, Zhenqi, et al.
Veröffentlicht: (2024) -
DiffAR: Denoising Diffusion Autoregressive Model for Raw Speech Waveform Generation
von: Benita, Roi, et al.
Veröffentlicht: (2023) -
Scaling Analysis of Interleaved Speech-Text Language Models
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)