Exploration of Perceptual Speech Features for Clinical Decision-Support in Mental Health Care
Fuente:
arXiv
Saved in:
| Main Authors: | Lyberatos, Vassilis, Dervakos, Edmund G., Adamidi, Eleni, Voulodimos, Athanasios, Stamou, Giorgos |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Perceptual Musical Features for Interpretable Audio Tagging
by: Lyberatos, Vassilis, et al.
Published: (2023)
by: Lyberatos, Vassilis, et al.
Published: (2023)
Exploring How Audio Effects Alter Emotion with Foundation Models
by: Katsis, Stelios, et al.
Published: (2025)
by: Katsis, Stelios, et al.
Published: (2025)
Go witheFlow: Real-time Emotion Driven Audio Effects Modulation
by: Dervakos, Edmund, et al.
Published: (2025)
by: Dervakos, Edmund, et al.
Published: (2025)
Semantic-Aware Interpretable Multimodal Music Auto-Tagging
by: Patakis, Andreas, et al.
Published: (2025)
by: Patakis, Andreas, et al.
Published: (2025)
CHORDONOMICON: A Dataset of 666,000 Songs and their Chord Progressions
by: Kantarelis, Spyridon, et al.
Published: (2024)
by: Kantarelis, Spyridon, et al.
Published: (2024)
MusicLIME: Explainable Multimodal Music Understanding
by: Sotirou, Theodoros, et al.
Published: (2024)
by: Sotirou, Theodoros, et al.
Published: (2024)
Semantic Prototypes: Enhancing Transparency Without Black Boxes
by: Menis-Mastromichalakis, Orfeas, et al.
Published: (2024)
by: Menis-Mastromichalakis, Orfeas, et al.
Published: (2024)
HalCECE: A Framework for Explainable Hallucination Detection through Conceptual Counterfactuals in Image Captioning
by: Lymperaiou, Maria, et al.
Published: (2025)
by: Lymperaiou, Maria, et al.
Published: (2025)
Finding My Voice: Generative Reconstruction of Disordered Speech for Automated Clinical Evaluation
by: Rosero, Karen, et al.
Published: (2025)
by: Rosero, Karen, et al.
Published: (2025)
SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
by: Zhang, Xueyao, et al.
Published: (2025)
by: Zhang, Xueyao, et al.
Published: (2025)
EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing
by: Sioros, Vassilis, et al.
Published: (2025)
by: Sioros, Vassilis, et al.
Published: (2025)
SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation
by: Liu, Ruohan, et al.
Published: (2026)
by: Liu, Ruohan, et al.
Published: (2026)
BERTtime Stories: Investigating the Role of Synthetic Story Data in Language Pre-training
by: Theodoropoulos, Nikitas, et al.
Published: (2024)
by: Theodoropoulos, Nikitas, et al.
Published: (2024)
Raon-Speech Technical Report
by: Kim, Beomsoo, et al.
Published: (2026)
by: Kim, Beomsoo, et al.
Published: (2026)
Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
by: Lin, Tzu-Quan, et al.
Published: (2025)
by: Lin, Tzu-Quan, et al.
Published: (2025)
EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
by: Zhang, Yuhao, et al.
Published: (2025)
by: Zhang, Yuhao, et al.
Published: (2025)
Density Adaptive Attention-based Speech Network: Enhancing Feature Understanding for Mental Health Disorders
by: Ioannides, Georgios, et al.
Published: (2024)
by: Ioannides, Georgios, et al.
Published: (2024)
V-CECE: Visual Counterfactual Explanations via Conceptual Edits
by: Spanos, Nikolaos, et al.
Published: (2025)
by: Spanos, Nikolaos, et al.
Published: (2025)
Harf-Speech: A Clinically Aligned Framework for Arabic Phoneme-Level Speech Assessment
by: Azad, Asif, et al.
Published: (2026)
by: Azad, Asif, et al.
Published: (2026)
Cross-Attention is Half Explanation in Speech-to-Text Models
by: Papi, Sara, et al.
Published: (2025)
by: Papi, Sara, et al.
Published: (2025)
Efficient Training for Cross-lingual Speech Language Models
by: Zhou, Yan, et al.
Published: (2026)
by: Zhou, Yan, et al.
Published: (2026)
Soundwave: Less is More for Speech-Text Alignment in LLMs
by: Zhang, Yuhao, et al.
Published: (2025)
by: Zhang, Yuhao, et al.
Published: (2025)
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
by: Papi, Sara, et al.
Published: (2025)
by: Papi, Sara, et al.
Published: (2025)
Conceptual Contrastive Edits in Textual and Vision-Language Retrieval
by: Lymperaiou, Maria, et al.
Published: (2025)
by: Lymperaiou, Maria, et al.
Published: (2025)
WESR: Scaling and Evaluating Word-level Event-Speech Recognition
by: Yang, Chenchen, et al.
Published: (2026)
by: Yang, Chenchen, et al.
Published: (2026)
Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models
by: Riera, Pablo, et al.
Published: (2026)
by: Riera, Pablo, et al.
Published: (2026)
Hybrid CNN-Transformer Architecture for Arabic Speech Emotion Recognition
by: Gheffari, Youcef Soufiane, et al.
Published: (2026)
by: Gheffari, Youcef Soufiane, et al.
Published: (2026)
Leveraging Audio and Text Modalities in Mental Health: A Study of LLMs Performance
by: Ali, Abdelrahman A., et al.
Published: (2024)
by: Ali, Abdelrahman A., et al.
Published: (2024)
Optimizing Speech Multi-View Feature Fusion through Conditional Computation
by: Shan, Weiqiao, et al.
Published: (2025)
by: Shan, Weiqiao, et al.
Published: (2025)
Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
by: Zhang, Xueyao, et al.
Published: (2025)
by: Zhang, Xueyao, et al.
Published: (2025)
Contextual Earnings-22: A Speech Recognition Benchmark with Custom Vocabulary in the Wild
by: Durmus, Berkin, et al.
Published: (2026)
by: Durmus, Berkin, et al.
Published: (2026)
Towards Fine-Grained Code-Switch Speech Translation with Semantic Space Alignment
by: Gao, Yan, et al.
Published: (2025)
by: Gao, Yan, et al.
Published: (2025)
MetaSICL: Adapting Audiroty LLM via Meta Speech In-Context Learning
by: Zheng, Haolong, et al.
Published: (2026)
by: Zheng, Haolong, et al.
Published: (2026)
AILS-NTUA at SemEval-2026 Task 12: Graph-Based Retrieval and Reflective Prompting for Abductive Event Reasoning
by: Karafyllis, Nikolas, et al.
Published: (2026)
by: Karafyllis, Nikolas, et al.
Published: (2026)
Pitfalls of Scale: Investigating the Inverse Task of Redefinition in Large Language Models
by: Stringli, Elena, et al.
Published: (2025)
by: Stringli, Elena, et al.
Published: (2025)
AILS-NTUA at SemEval-2026 Task 8: Evaluating Multi-Turn RAG Conversations
by: Athanasiou, Dimosthenis, et al.
Published: (2026)
by: Athanasiou, Dimosthenis, et al.
Published: (2026)
AILS-NTUA at SemEval-2025 Task 8: Language-to-Code prompting and Error Fixing for Tabular Question Answering
by: Evangelatos, Andreas, et al.
Published: (2025)
by: Evangelatos, Andreas, et al.
Published: (2025)
Data-efficient Targeted Token-level Preference Optimization for LLM-based Text-to-Speech
by: Kotoge, Rikuto, et al.
Published: (2025)
by: Kotoge, Rikuto, et al.
Published: (2025)
FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian
by: Papi, Sara, et al.
Published: (2025)
by: Papi, Sara, et al.
Published: (2025)
Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided Optimization
by: Shi, Jiacheng, et al.
Published: (2025)
by: Shi, Jiacheng, et al.
Published: (2025)
Similar Items
-
Perceptual Musical Features for Interpretable Audio Tagging
by: Lyberatos, Vassilis, et al.
Published: (2023) -
Exploring How Audio Effects Alter Emotion with Foundation Models
by: Katsis, Stelios, et al.
Published: (2025) -
Go witheFlow: Real-time Emotion Driven Audio Effects Modulation
by: Dervakos, Edmund, et al.
Published: (2025) -
Semantic-Aware Interpretable Multimodal Music Auto-Tagging
by: Patakis, Andreas, et al.
Published: (2025) -
CHORDONOMICON: A Dataset of 666,000 Songs and their Chord Progressions
by: Kantarelis, Spyridon, et al.
Published: (2024)