Foundation Model-based Evaluation of Neuropsychiatric Disorders: A Lifespan-Inclusive, Multi-Modal, and Multi-Lingual Study
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dong, Zhongren, Guo, Haotian, Xu, Weixiang, Zhao, Huan, Zhang, Zixing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Re-Parameterization of Lightweight Transformer for On-Device Speech Emotion Recognition
von: Zhang, Zixing, et al.
Veröffentlicht: (2024)
von: Zhang, Zixing, et al.
Veröffentlicht: (2024)
HAFFormer: A Hierarchical Attention-Free Framework for Alzheimer's Disease Detection From Spontaneous Speech
von: Dong, Zhongren, et al.
Veröffentlicht: (2024)
von: Dong, Zhongren, et al.
Veröffentlicht: (2024)
ParaLBench: A Large-Scale Benchmark for Computational Paralinguistics over Acoustic Foundation Models
von: Zhang, Zixing, et al.
Veröffentlicht: (2024)
von: Zhang, Zixing, et al.
Veröffentlicht: (2024)
Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition and Translation
von: Zhao, Qiuming, et al.
Veröffentlicht: (2025)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2025)
SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs
von: Dong, Zhongren, et al.
Veröffentlicht: (2025)
von: Dong, Zhongren, et al.
Veröffentlicht: (2025)
Improving X-Codec-2.0 for Multi-Lingual Speech: 25 Hz Latent Rate and 24 kHz Sampling
von: Zolkepli, Husein
Veröffentlicht: (2026)
von: Zolkepli, Husein
Veröffentlicht: (2026)
Leveraging Parameter-Efficient Transfer Learning for Multi-Lingual Text-to-Speech Adaptation
von: Li, Yingting, et al.
Veröffentlicht: (2024)
von: Li, Yingting, et al.
Veröffentlicht: (2024)
Cross-Lingual Multi-Granularity Framework for Interpretable Parkinson's Disease Diagnosis from Speech
von: Tougui, Ilias, et al.
Veröffentlicht: (2025)
von: Tougui, Ilias, et al.
Veröffentlicht: (2025)
POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation
von: Li, Xuanchen, et al.
Veröffentlicht: (2025)
von: Li, Xuanchen, et al.
Veröffentlicht: (2025)
Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment
von: Wang, Xuechen, et al.
Veröffentlicht: (2024)
von: Wang, Xuechen, et al.
Veröffentlicht: (2024)
Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling
von: Cheng, Yifan, et al.
Veröffentlicht: (2025)
von: Cheng, Yifan, et al.
Veröffentlicht: (2025)
CLiFT-ASR: A Cross-Lingual Fine-Tuning Framework for Low-Resource Taiwanese Hokkien Speech Recognition
von: Sung, Hung-Yang, et al.
Veröffentlicht: (2025)
von: Sung, Hung-Yang, et al.
Veröffentlicht: (2025)
ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis
von: He, Xiangheng, et al.
Veröffentlicht: (2024)
von: He, Xiangheng, et al.
Veröffentlicht: (2024)
Multi-Modal Automatic Prosody Annotation with Contrastive Pretraining of SSWP
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2023)
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2023)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
von: Gosai, Advait, et al.
Veröffentlicht: (2025)
von: Gosai, Advait, et al.
Veröffentlicht: (2025)
Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models
von: Lin, Yu-Xiang, et al.
Veröffentlicht: (2025)
von: Lin, Yu-Xiang, et al.
Veröffentlicht: (2025)
Data-Centric Improvements for Enhancing Multi-Modal Understanding in Spoken Conversation Modeling
von: Chen, Maximillian, et al.
Veröffentlicht: (2024)
von: Chen, Maximillian, et al.
Veröffentlicht: (2024)
SSR: Alignment-Aware Modality Connector for Speech Language Models
von: Tan, Weiting, et al.
Veröffentlicht: (2024)
von: Tan, Weiting, et al.
Veröffentlicht: (2024)
BRACE: A Benchmark for Robust Audio Caption Quality Evaluation
von: Guo, Tianyu, et al.
Veröffentlicht: (2025)
von: Guo, Tianyu, et al.
Veröffentlicht: (2025)
Equipping LLM with Directional Multi-Talker Speech Understanding Capabilities
von: Lin, Ju, et al.
Veröffentlicht: (2026)
von: Lin, Ju, et al.
Veröffentlicht: (2026)
DELULU: Discriminative Embedding Learning Using Latent Units for Speaker-Aware Self-Trained Speech Foundational Model
von: Baali, Massa, et al.
Veröffentlicht: (2025)
von: Baali, Massa, et al.
Veröffentlicht: (2025)
Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control
von: Yamamoto, Ryuichi, et al.
Veröffentlicht: (2024)
von: Yamamoto, Ryuichi, et al.
Veröffentlicht: (2024)
FCPE: A Fast Context-based Pitch Estimation Model
von: Luo, Yuxin, et al.
Veröffentlicht: (2025)
von: Luo, Yuxin, et al.
Veröffentlicht: (2025)
Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play
von: Shi, Yemin, et al.
Veröffentlicht: (2025)
von: Shi, Yemin, et al.
Veröffentlicht: (2025)
M-CIF: Multi-Scale Alignment For CIF-Based Non-Autoregressive ASR
von: Mao, Ruixiang, et al.
Veröffentlicht: (2025)
von: Mao, Ruixiang, et al.
Veröffentlicht: (2025)
Finding My Voice: Generative Reconstruction of Disordered Speech for Automated Clinical Evaluation
von: Rosero, Karen, et al.
Veröffentlicht: (2025)
von: Rosero, Karen, et al.
Veröffentlicht: (2025)
Robust Audio-Text Retrieval via Cross-Modal Attention and Hybrid Loss
von: Liu, Meizhu, et al.
Veröffentlicht: (2026)
von: Liu, Meizhu, et al.
Veröffentlicht: (2026)
DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2026)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2026)
Large Language Model Should Understand Pinyin for Chinese ASR Error Correction
von: Li, Yuang, et al.
Veröffentlicht: (2024)
von: Li, Yuang, et al.
Veröffentlicht: (2024)
Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
VocalNet-M2: Advancing Low-Latency Spoken Language Modeling via Integrated Multi-Codebook Tokenization and Multi-Token Prediction
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
Multi-Modal Retrieval For Large Language Model Based Speech Recognition
von: Kolehmainen, Jari, et al.
Veröffentlicht: (2024)
von: Kolehmainen, Jari, et al.
Veröffentlicht: (2024)
DementiaBank-Emotion: A Multi-Rater Emotion Annotation Corpus for Alzheimer's Disease Speech (Version 1.0)
von: Jeong, Cheonkam, et al.
Veröffentlicht: (2026)
von: Jeong, Cheonkam, et al.
Veröffentlicht: (2026)
MultiPA: A Multi-task Speech Pronunciation Assessment Model for Open Response Scenarios
von: Chen, Yu-Wen, et al.
Veröffentlicht: (2023)
von: Chen, Yu-Wen, et al.
Veröffentlicht: (2023)
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Re-Parameterization of Lightweight Transformer for On-Device Speech Emotion Recognition
von: Zhang, Zixing, et al.
Veröffentlicht: (2024) -
HAFFormer: A Hierarchical Attention-Free Framework for Alzheimer's Disease Detection From Spontaneous Speech
von: Dong, Zhongren, et al.
Veröffentlicht: (2024) -
ParaLBench: A Large-Scale Benchmark for Computational Paralinguistics over Acoustic Foundation Models
von: Zhang, Zixing, et al.
Veröffentlicht: (2024) -
Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition and Translation
von: Zhao, Qiuming, et al.
Veröffentlicht: (2025) -
SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs
von: Dong, Zhongren, et al.
Veröffentlicht: (2025)