Emotional Vietnamese Speech-Based Depression Diagnosis Using Dynamic Attention Mechanism
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | D., Quang-Anh N., Ha, Manh-Hung, Dinh, Thai Kim, Pham, Minh-Duc, Van, Ninh Nguyen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TSPC: A Two-Stage Phoneme-Centric Architecture for code-switching Vietnamese-English Speech Recognition
von: Anh, Tran Nguyen, et al.
Veröffentlicht: (2025)
von: Anh, Tran Nguyen, et al.
Veröffentlicht: (2025)
Improving Streaming Speech Recognition With Time-Shifted Contextual Attention And Dynamic Right Context Masking
von: Le, Khanh, et al.
Veröffentlicht: (2025)
von: Le, Khanh, et al.
Veröffentlicht: (2025)
Zero-Shot Text-to-Speech for Vietnamese
von: Vu, Thi, et al.
Veröffentlicht: (2025)
von: Vu, Thi, et al.
Veröffentlicht: (2025)
BERSting at the Screams: A Benchmark for Distanced, Emotional and Shouted Speech Recognition
von: Tuttösí, Paige, et al.
Veröffentlicht: (2025)
von: Tuttösí, Paige, et al.
Veröffentlicht: (2025)
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
von: Pham, The Hieu, et al.
Veröffentlicht: (2025)
von: Pham, The Hieu, et al.
Veröffentlicht: (2025)
Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof Detection
von: Zhao, Junchuan, et al.
Veröffentlicht: (2026)
von: Zhao, Junchuan, et al.
Veröffentlicht: (2026)
MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
von: Le-Duc, Khai, et al.
Veröffentlicht: (2024)
von: Le-Duc, Khai, et al.
Veröffentlicht: (2024)
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
von: Hu, Cheng-Hung, et al.
Veröffentlicht: (2025)
von: Hu, Cheng-Hung, et al.
Veröffentlicht: (2025)
Dataset-Distillation Generative Model for Speech Emotion Recognition
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2024)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2024)
EMO-SUPERB: An In-depth Look at Speech Emotion Recognition
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Adaptive Speech Emotion Representation Learning Based On Dynamic Graph
von: Gao, Yingxue, et al.
Veröffentlicht: (2024)
von: Gao, Yingxue, et al.
Veröffentlicht: (2024)
Speechless: Speech Instruction Training Without Speech for Low Resource Languages
von: Dao, Alan, et al.
Veröffentlicht: (2025)
von: Dao, Alan, et al.
Veröffentlicht: (2025)
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling
von: Pham, Long-Khanh, et al.
Veröffentlicht: (2025)
von: Pham, Long-Khanh, et al.
Veröffentlicht: (2025)
Double Multi-Head Attention Multimodal System for Odyssey 2024 Speech Emotion Recognition Challenge
von: Costa, Federico, et al.
Veröffentlicht: (2024)
von: Costa, Federico, et al.
Veröffentlicht: (2024)
Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model
von: Ueda, Lucas, et al.
Veröffentlicht: (2025)
von: Ueda, Lucas, et al.
Veröffentlicht: (2025)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
Metadata-Enhanced Speech Emotion Recognition: Augmented Residual Integration and Co-Attention in Two-Stage Fine-Tuning
von: Wan, Zixiang, et al.
Veröffentlicht: (2024)
von: Wan, Zixiang, et al.
Veröffentlicht: (2024)
Textless and Non-Parallel Speech-to-Speech Emotion Style Transfer
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
Speech Emotion Recognition with ASR Integration
von: Li, Yuanchao
Veröffentlicht: (2026)
von: Li, Yuanchao
Veröffentlicht: (2026)
Attention-weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition
von: Yang, Qingran, et al.
Veröffentlicht: (2026)
von: Yang, Qingran, et al.
Veröffentlicht: (2026)
VoxVietnam: a Large-Scale Multi-Genre Dataset for Vietnamese Speaker Recognition
von: Vu, Hoang Long, et al.
Veröffentlicht: (2024)
von: Vu, Hoang Long, et al.
Veröffentlicht: (2024)
Searching for Effective Preprocessing Method and CNN-based Architecture with Efficient Channel Attention on Speech Emotion Recognition
von: Kim, Byunggun, et al.
Veröffentlicht: (2024)
von: Kim, Byunggun, et al.
Veröffentlicht: (2024)
Hierarchical Control of Emotion Rendering in Speech Synthesis
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis
von: Tang, Haobin, et al.
Veröffentlicht: (2024)
von: Tang, Haobin, et al.
Veröffentlicht: (2024)
Temporal-Channel Modeling in Multi-head Self-Attention for Synthetic Speech Detection
von: Truong, Duc-Tuan, et al.
Veröffentlicht: (2024)
von: Truong, Duc-Tuan, et al.
Veröffentlicht: (2024)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
Fine-Grained Quantitative Emotion Editing for Speech Generation
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
THAI Speech Emotion Recognition (THAI-SER) corpus
von: Wongpithayadisai, Jilamika, et al.
Veröffentlicht: (2025)
von: Wongpithayadisai, Jilamika, et al.
Veröffentlicht: (2025)
Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
von: Sun, Haoqin, et al.
Veröffentlicht: (2024)
von: Sun, Haoqin, et al.
Veröffentlicht: (2024)
Naturalness-Aware Curriculum Learning with Dynamic Temperature for Speech Deepfake Detection
von: Kim, Taewoo, et al.
Veröffentlicht: (2025)
von: Kim, Taewoo, et al.
Veröffentlicht: (2025)
Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
von: Zhao, Yan, et al.
Veröffentlicht: (2024)
von: Zhao, Yan, et al.
Veröffentlicht: (2024)
End-to-End Target Speaker Speech Recognition Using Context-Aware Attention Mechanisms for Challenging Enrollment Scenario
von: Ghane, Mohsen, et al.
Veröffentlicht: (2025)
von: Ghane, Mohsen, et al.
Veröffentlicht: (2025)
PARROT: Synergizing Mamba and Attention-based SSL Pre-Trained Models via Parallel Branch Hadamard Optimal Transport for Speech Emotion Recognition
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
Efficient Long Speech Sequence Modelling for Time-Domain Depression Level Estimation
von: Li, Shuanglin, et al.
Veröffentlicht: (2025)
von: Li, Shuanglin, et al.
Veröffentlicht: (2025)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024)
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024)
Semantic-Emotional Resonance Embedding: A Semi-Supervised Paradigm for Cross-Lingual Speech Emotion Recognition
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
von: Pham, Lam, et al.
Veröffentlicht: (2024)
von: Pham, Lam, et al.
Veröffentlicht: (2024)
Mamba in Speech: Towards an Alternative to Self-Attention
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
PCQ: Emotion Recognition in Speech via Progressive Channel Querying
von: Wang, Xincheng, et al.
Veröffentlicht: (2024)
von: Wang, Xincheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TSPC: A Two-Stage Phoneme-Centric Architecture for code-switching Vietnamese-English Speech Recognition
von: Anh, Tran Nguyen, et al.
Veröffentlicht: (2025) -
Improving Streaming Speech Recognition With Time-Shifted Contextual Attention And Dynamic Right Context Masking
von: Le, Khanh, et al.
Veröffentlicht: (2025) -
Zero-Shot Text-to-Speech for Vietnamese
von: Vu, Thi, et al.
Veröffentlicht: (2025) -
BERSting at the Screams: A Benchmark for Distanced, Emotional and Shouted Speech Recognition
von: Tuttösí, Paige, et al.
Veröffentlicht: (2025) -
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
von: Pham, The Hieu, et al.
Veröffentlicht: (2025)