EmoSURA: Towards Accurate Evaluation of Detailed and Long-Context Emotional Speech Captions
Fuente:
arXiv
Salvato in:
| Autori principali: | Jing, Xin, Triantafyllopoulos, Andreas, Wang, Jiadong, Amiriparian, Shahin, Luo, Jun, Schuller, Björn |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SmoothCLAP: Soft-Target Enhanced Contrastive Language\--Audio Pretraining for Affective Computing
di: Jing, Xin, et al.
Pubblicazione: (2026)
di: Jing, Xin, et al.
Pubblicazione: (2026)
MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge
di: Jing, Xin, et al.
Pubblicazione: (2025)
di: Jing, Xin, et al.
Pubblicazione: (2025)
ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks
di: Jing, Xin, et al.
Pubblicazione: (2024)
di: Jing, Xin, et al.
Pubblicazione: (2024)
Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
di: Jing, Xin, et al.
Pubblicazione: (2024)
di: Jing, Xin, et al.
Pubblicazione: (2024)
Non-Verbal Vocalisations and their Challenges: Emotion, Privacy, Sparseness, and Real Life
di: Batliner, Anton, et al.
Pubblicazione: (2025)
di: Batliner, Anton, et al.
Pubblicazione: (2025)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
di: Qi, Tianhua, et al.
Pubblicazione: (2026)
di: Qi, Tianhua, et al.
Pubblicazione: (2026)
Abusive Speech Detection in Indic Languages Using Acoustic Features
di: Spiesberger, Anika A., et al.
Pubblicazione: (2024)
di: Spiesberger, Anika A., et al.
Pubblicazione: (2024)
Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2025)
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2025)
Towards Multimodal Prediction of Spontaneous Humour: A Novel Dataset and First Results
di: Christ, Lukas, et al.
Pubblicazione: (2022)
di: Christ, Lukas, et al.
Pubblicazione: (2022)
Audio-based Step-count Estimation for Running -- Windowing and Neural Network Baselines
di: Wagner, Philipp, et al.
Pubblicazione: (2024)
di: Wagner, Philipp, et al.
Pubblicazione: (2024)
EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
di: Bian, Weizhen, et al.
Pubblicazione: (2024)
di: Bian, Weizhen, et al.
Pubblicazione: (2024)
EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses
di: Xu, Shuhao, et al.
Pubblicazione: (2026)
di: Xu, Shuhao, et al.
Pubblicazione: (2026)
Non-Invasive Suicide Risk Prediction Through Speech Analysis
di: Amiriparian, Shahin, et al.
Pubblicazione: (2024)
di: Amiriparian, Shahin, et al.
Pubblicazione: (2024)
Exploring Gender-Specific Speech Patterns in Automatic Suicide Risk Assessment
di: Gerczuk, Maurice, et al.
Pubblicazione: (2024)
di: Gerczuk, Maurice, et al.
Pubblicazione: (2024)
A conceptual framework for learning to listen by reward: Curiosity-driven search for novel sources
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2026)
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2026)
An automatic analysis of ultrasound vocalisations for the prediction of interaction context in captive Egyptian fruit bats
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2024)
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2024)
autrainer: A Modular and Extensible Deep Learning Toolkit for Computer Audition Tasks
di: Rampp, Simon, et al.
Pubblicazione: (2024)
di: Rampp, Simon, et al.
Pubblicazione: (2024)
EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
di: Cong, Gaoxiang, et al.
Pubblicazione: (2024)
di: Cong, Gaoxiang, et al.
Pubblicazione: (2024)
Bringing the Discussion of Minima Sharpness to the Audio Domain: a Filter-Normalised Evaluation for Acoustic Scene Classification
di: Milling, Manuel, et al.
Pubblicazione: (2023)
di: Milling, Manuel, et al.
Pubblicazione: (2023)
Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach
di: Sankey-Olsen, Cuno, et al.
Pubblicazione: (2025)
di: Sankey-Olsen, Cuno, et al.
Pubblicazione: (2025)
Computer Audition: From Task-Specific Machine Learning to Foundation Models
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2024)
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2024)
Audio Enhancement for Computer Audition -- An Iterative Training Paradigm Using Sample Importance
di: Milling, Manuel, et al.
Pubblicazione: (2024)
di: Milling, Manuel, et al.
Pubblicazione: (2024)
This Paper Had the Smartest Reviewers -- Flattery Detection Utilising an Audio-Textual Transformer-Based Approach
di: Christ, Lukas, et al.
Pubblicazione: (2024)
di: Christ, Lukas, et al.
Pubblicazione: (2024)
DFingerNet: Noise-Adaptive Speech Enhancement for Hearing Aids
di: Tsangko, Iosif, et al.
Pubblicazione: (2025)
di: Tsangko, Iosif, et al.
Pubblicazione: (2025)
Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
di: Zhao, Yan, et al.
Pubblicazione: (2024)
di: Zhao, Yan, et al.
Pubblicazione: (2024)
Enhancing Speech Emotion Recognition Through Differentiable Architecture Search
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2023)
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2023)
Can Large Language Models Aid in Annotating Speech Emotional Data? Uncovering New Frontiers
di: Latif, Siddique, et al.
Pubblicazione: (2023)
di: Latif, Siddique, et al.
Pubblicazione: (2023)
Domain Adapting Deep Reinforcement Learning for Real-world Speech Emotion Recognition
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2022)
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2022)
Testing Correctness, Fairness, and Robustness of Speech Emotion Recognition Models
di: Derington, Anna, et al.
Pubblicazione: (2023)
di: Derington, Anna, et al.
Pubblicazione: (2023)
Exploring Meta Information for Audio-based Zero-shot Bird Classification
di: Gebhard, Alexander, et al.
Pubblicazione: (2023)
di: Gebhard, Alexander, et al.
Pubblicazione: (2023)
The Affective Bridge: Preserving Speech Representations while Enhancing Deepfake Detection vian emotional Constraints
di: Li, Yupei, et al.
Pubblicazione: (2025)
di: Li, Yupei, et al.
Pubblicazione: (2025)
EmoHRNet: High-Resolution Neural Network Based Speech Emotion Recognition
di: Muppidi, Akshay, et al.
Pubblicazione: (2025)
di: Muppidi, Akshay, et al.
Pubblicazione: (2025)
Representation Learning with Parameterised Quantum Circuits for Advancing Speech Emotion Recognition
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2025)
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2025)
Are you sure? Analysing Uncertainty Quantification Approaches for Real-world Speech Emotion Recognition
di: Schrüfer, Oliver, et al.
Pubblicazione: (2024)
di: Schrüfer, Oliver, et al.
Pubblicazione: (2024)
Expressivity and Speech Synthesis
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2024)
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2024)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
di: Zhang, Hezhao, et al.
Pubblicazione: (2026)
di: Zhang, Hezhao, et al.
Pubblicazione: (2026)
EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model
di: Yang, Yiqing, et al.
Pubblicazione: (2025)
di: Yang, Yiqing, et al.
Pubblicazione: (2025)
Improving Speaker-independent Speech Emotion Recognition Using Dynamic Joint Distribution Adaptation
di: Lu, Cheng, et al.
Pubblicazione: (2024)
di: Lu, Cheng, et al.
Pubblicazione: (2024)
BLSP-Emo: Towards Empathetic Large Speech-Language Models
di: Wang, Chen, et al.
Pubblicazione: (2024)
di: Wang, Chen, et al.
Pubblicazione: (2024)
STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition
di: Chang, Yi, et al.
Pubblicazione: (2024)
di: Chang, Yi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SmoothCLAP: Soft-Target Enhanced Contrastive Language\--Audio Pretraining for Affective Computing
di: Jing, Xin, et al.
Pubblicazione: (2026) -
MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge
di: Jing, Xin, et al.
Pubblicazione: (2025) -
ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks
di: Jing, Xin, et al.
Pubblicazione: (2024) -
Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
di: Jing, Xin, et al.
Pubblicazione: (2024) -
Non-Verbal Vocalisations and their Challenges: Emotion, Privacy, Sparseness, and Real Life
di: Batliner, Anton, et al.
Pubblicazione: (2025)