Hierarchical Self-Supervised Representation Learning for Depression Detection from Speech
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yuxin, Chng, Eng Siong, Guan, Cuntai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs
von: Yuhang, Yang, et al.
Veröffentlicht: (2024)
von: Yuhang, Yang, et al.
Veröffentlicht: (2024)
Temporal-Channel Modeling in Multi-head Self-Attention for Synthetic Speech Detection
von: Truong, Duc-Tuan, et al.
Veröffentlicht: (2024)
von: Truong, Duc-Tuan, et al.
Veröffentlicht: (2024)
Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation Models
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Continual Learning Optimizations for Auto-regressive Decoder of Multilingual ASR systems
von: Kwok, Chin Yuen, et al.
Veröffentlicht: (2024)
von: Kwok, Chin Yuen, et al.
Veröffentlicht: (2024)
GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models
von: Deng, Yayue, et al.
Veröffentlicht: (2025)
von: Deng, Yayue, et al.
Veröffentlicht: (2025)
Speech Enhancement Using Continuous Embeddings of Neural Audio Codec
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
von: Sun, Yujia, et al.
Veröffentlicht: (2024)
von: Sun, Yujia, et al.
Veröffentlicht: (2024)
Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning
von: He, Haorui, et al.
Veröffentlicht: (2024)
von: He, Haorui, et al.
Veröffentlicht: (2024)
ASDA: Audio Spectrogram Differential Attention Mechanism for Self-Supervised Representation Learning
von: Wang, Junyu, et al.
Veröffentlicht: (2025)
von: Wang, Junyu, et al.
Veröffentlicht: (2025)
Training-Free Intelligibility-Guided Observation Addition for Noisy ASR
von: Li, Haoyang, et al.
Veröffentlicht: (2026)
von: Li, Haoyang, et al.
Veröffentlicht: (2026)
Exploring Self-Supervised Multi-view Contrastive Learning for Speech Emotion Recognition with Limited Annotations
von: Khaertdinov, Bulat, et al.
Veröffentlicht: (2024)
von: Khaertdinov, Bulat, et al.
Veröffentlicht: (2024)
Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
von: Chen, Chen, et al.
Veröffentlicht: (2025)
von: Chen, Chen, et al.
Veröffentlicht: (2025)
Self-supervised Speech Models for Word-Level Stuttered Speech Detection
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
Quantizer-Aware Hierarchical Neural Codec Modeling for Speech Deepfake Detection
von: Wu, Jinyang, et al.
Veröffentlicht: (2026)
von: Wu, Jinyang, et al.
Veröffentlicht: (2026)
Towards Robust Speech Representation Learning for Thousands of Languages
von: Chen, William, et al.
Veröffentlicht: (2024)
von: Chen, William, et al.
Veröffentlicht: (2024)
Large Language Models are Efficient Learners of Noise-Robust Speech Recognition
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
Jointly Fine-Tuning "BERT-like" Self Supervised Models to Improve Multimodal Speech Emotion Recognition
von: Siriwardhana, Shamane, et al.
Veröffentlicht: (2020)
von: Siriwardhana, Shamane, et al.
Veröffentlicht: (2020)
Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
Speechless: Speech Instruction Training Without Speech for Low Resource Languages
von: Dao, Alan, et al.
Veröffentlicht: (2025)
von: Dao, Alan, et al.
Veröffentlicht: (2025)
Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson's Disease
von: Hernandez, Abner, et al.
Veröffentlicht: (2026)
von: Hernandez, Abner, et al.
Veröffentlicht: (2026)
LeBenchmark 2.0: a Standardized, Replicable and Enhanced Framework for Self-supervised Representations of French Speech
von: Parcollet, Titouan, et al.
Veröffentlicht: (2023)
von: Parcollet, Titouan, et al.
Veröffentlicht: (2023)
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
Virtual Speech Therapist: A Clinician-in-the-Loop AI Speech Therapy Agent for Personalized and Supervised Therapy
von: Sheikh, Shakeel, et al.
Veröffentlicht: (2026)
von: Sheikh, Shakeel, et al.
Veröffentlicht: (2026)
VQ-CTAP: Cross-Modal Fine-Grained Sequence Representation Learning for Speech Processing
von: Qiang, Chunyu, et al.
Veröffentlicht: (2024)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2024)
Wav2code: Restore Clean Speech Representations via Codebook Lookup for Noise-Robust ASR
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
Noise-Aware Speech Separation with Contrastive Learning
von: Zhang, Zizheng, et al.
Veröffentlicht: (2023)
von: Zhang, Zizheng, et al.
Veröffentlicht: (2023)
DM-Codec: Distilling Multimodal Representations for Speech Tokenization
von: Ahasan, Md Mubtasim, et al.
Veröffentlicht: (2024)
von: Ahasan, Md Mubtasim, et al.
Veröffentlicht: (2024)
StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024)
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024)
UniArray: Unified Spectral-Spatial Modeling for Array-Geometry-Agnostic Speech Separation
von: Chen, Weiguang, et al.
Veröffentlicht: (2025)
von: Chen, Weiguang, et al.
Veröffentlicht: (2025)
WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning
von: Rao, Rajath, et al.
Veröffentlicht: (2025)
von: Rao, Rajath, et al.
Veröffentlicht: (2025)
An Investigation Into Explainable Audio Hate Speech Detection
von: An, Jinmyeong, et al.
Veröffentlicht: (2024)
von: An, Jinmyeong, et al.
Veröffentlicht: (2024)
Beyond Hard Sharing: Efficient Multi-Task Speech-to-Text Modeling with Supervised Mixture of Experts
von: Jin, Hojun, et al.
Veröffentlicht: (2025)
von: Jin, Hojun, et al.
Veröffentlicht: (2025)
A Benchmark for Early-stage Parkinson's Disease Detection from Speech
von: Zhong, Terry Yi, et al.
Veröffentlicht: (2026)
von: Zhong, Terry Yi, et al.
Veröffentlicht: (2026)
ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models
von: Hsiao, Chi-Yuan, et al.
Veröffentlicht: (2026)
von: Hsiao, Chi-Yuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs
von: Yuhang, Yang, et al.
Veröffentlicht: (2024) -
Temporal-Channel Modeling in Multi-head Self-Attention for Synthetic Speech Detection
von: Truong, Duc-Tuan, et al.
Veröffentlicht: (2024) -
Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models
von: Hu, Yuchen, et al.
Veröffentlicht: (2024) -
Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation Models
von: Hu, Yuchen, et al.
Veröffentlicht: (2024) -
Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)