A Novel Markovian Framework for Integrating Absolute and Relative Ordinal Emotion Information
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Jingyao, Dang, Ting, Sethu, Vidhyasaharan, Ambikairajah, Eliathamby |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2021
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Characterization of Speech Similarity Between Australian Aboriginal and High-Resource Languages: A Case Study on Dharawal
von: Dang, Ting, et al.
Veröffentlicht: (2025)
von: Dang, Ting, et al.
Veröffentlicht: (2025)
What is Learnt by the LEArnable Front-end (LEAF)? Adapting Per-Channel Energy Normalisation (PCEN) to Noisy Conditions
von: Meng, Hanyu, et al.
Veröffentlicht: (2024)
von: Meng, Hanyu, et al.
Veröffentlicht: (2024)
Should Audio Front-ends be Adaptive? Comparing Learnable and Adaptive Front-ends
von: Zhang, Qiquan, et al.
Veröffentlicht: (2025)
von: Zhang, Qiquan, et al.
Veröffentlicht: (2025)
Binaural Selective Attention Model for Target Speaker Extraction
von: Meng, Hanyu, et al.
Veröffentlicht: (2024)
von: Meng, Hanyu, et al.
Veröffentlicht: (2024)
Adaptive Per-Channel Energy Normalization Front-end for Robust Audio Signal Processing
von: Meng, Hanyu, et al.
Veröffentlicht: (2025)
von: Meng, Hanyu, et al.
Veröffentlicht: (2025)
Blind Estimation of Sub-band Acoustic Parameters from Ambisonics Recordings using Spectro-Spatial Covariance Features
von: Meng, Hanyu, et al.
Veröffentlicht: (2024)
von: Meng, Hanyu, et al.
Veröffentlicht: (2024)
Dual-Constrained Dynamical Neural ODEs for Ambiguity-aware Continuous Emotion Prediction
von: Wu, Jingyao, et al.
Veröffentlicht: (2024)
von: Wu, Jingyao, et al.
Veröffentlicht: (2024)
A Joint Spectro-Temporal Relational Thinking Based Acoustic Modeling Framework
von: Nan, Zheng, et al.
Veröffentlicht: (2024)
von: Nan, Zheng, et al.
Veröffentlicht: (2024)
Decoding Ambiguous Emotions with Test-Time Scaling in Audio-Language Models
von: Jia, Hong, et al.
Veröffentlicht: (2026)
von: Jia, Hong, et al.
Veröffentlicht: (2026)
Exploring Length Generalization For Transformer-based Speech Enhancement
von: Zhang, Qiquan, et al.
Veröffentlicht: (2025)
von: Zhang, Qiquan, et al.
Veröffentlicht: (2025)
An Exploration of Length Generalization in Transformer-Based Speech Enhancement
von: Zhang, Qiquan, et al.
Veröffentlicht: (2024)
von: Zhang, Qiquan, et al.
Veröffentlicht: (2024)
Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
von: Adelson, Trevor, et al.
Veröffentlicht: (2026)
von: Adelson, Trevor, et al.
Veröffentlicht: (2026)
Continual Adaptation for Pacific Indigenous Speech Recognition
von: Xiao, Yang, et al.
Veröffentlicht: (2026)
von: Xiao, Yang, et al.
Veröffentlicht: (2026)
Why Can't They Remember? Uncovering Representation and Retrieval Bottlenecks in Multi-Turn Acoustic Memory
von: Xiao, Yang, et al.
Veröffentlicht: (2026)
von: Xiao, Yang, et al.
Veröffentlicht: (2026)
Disentangling Reasoning in Large Audio-Language Models for Ambiguous Emotion Prediction
von: Yu, Xiaofeng, et al.
Veröffentlicht: (2026)
von: Yu, Xiaofeng, et al.
Veröffentlicht: (2026)
Moonbeam: A MIDI Foundation Model Using Both Absolute and Relative Music Attributes
von: Guo, Zixun, et al.
Veröffentlicht: (2025)
von: Guo, Zixun, et al.
Veröffentlicht: (2025)
An Empirical Study on the Impact of Positional Encoding in Transformer-based Monaural Speech Enhancement
von: Zhang, Qiquan, et al.
Veröffentlicht: (2024)
von: Zhang, Qiquan, et al.
Veröffentlicht: (2024)
Scaling Auditory Cognition via Test-Time Compute in Audio Language Models
von: Dang, Ting, et al.
Veröffentlicht: (2025)
von: Dang, Ting, et al.
Veröffentlicht: (2025)
Zero-Shot Text-to-Speech as Golden Speech Generator: A Systematic Framework and its Applicability in Automatic Pronunciation Assessment
von: Lo, Tien-Hong, et al.
Veröffentlicht: (2024)
von: Lo, Tien-Hong, et al.
Veröffentlicht: (2024)
Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement
von: Yang, Jianing, et al.
Veröffentlicht: (2025)
von: Yang, Jianing, et al.
Veröffentlicht: (2025)
Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models
von: Zhang, Wenda, et al.
Veröffentlicht: (2026)
von: Zhang, Wenda, et al.
Veröffentlicht: (2026)
SPUR: A Plug-and-Play Framework for Integrating Spatial Audio Understanding and Reasoning into Large Audio-Language Models
von: Sakshi, S, et al.
Veröffentlicht: (2025)
von: Sakshi, S, et al.
Veröffentlicht: (2025)
Qieemo: Speech Is All You Need in the Emotion Recognition in Conversations
von: Chen, Jinming, et al.
Veröffentlicht: (2025)
von: Chen, Jinming, et al.
Veröffentlicht: (2025)
Sync-TVA: A Graph-Attention Framework for Multimodal Emotion Recognition with Cross-Modal Fusion
von: Deng, Zeyu, et al.
Veröffentlicht: (2025)
von: Deng, Zeyu, et al.
Veröffentlicht: (2025)
Color-based Emotion Representation for Speech Emotion Recognition
von: Nagase, Ryotaro, et al.
Veröffentlicht: (2026)
von: Nagase, Ryotaro, et al.
Veröffentlicht: (2026)
Mamba in Speech: Towards an Alternative to Self-Attention
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
Investigating Acoustic-Textual Emotional Inconsistency Information for Automatic Depression Detection
von: Su, Rongfeng, et al.
Veröffentlicht: (2024)
von: Su, Rongfeng, et al.
Veröffentlicht: (2024)
Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
Dual-branch Graph Domain Adaptation for Cross-scenario Multi-modal Emotion Recognition
von: Shou, Yuntao, et al.
Veröffentlicht: (2026)
von: Shou, Yuntao, et al.
Veröffentlicht: (2026)
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
Joint Learning using Mixture-of-Expert-Based Representation for Speech Enhancement and Robust Emotion Recognition
von: Tzeng, Jing-Tong, et al.
Veröffentlicht: (2025)
von: Tzeng, Jing-Tong, et al.
Veröffentlicht: (2025)
Spatiotemporal Emotional Synchrony in Dyadic Interactions: The Role of Speech Conditions in Facial and Vocal Affective Alignment
von: Herbuela, Von Ralph Dane Marquez, et al.
Veröffentlicht: (2025)
von: Herbuela, Von Ralph Dane Marquez, et al.
Veröffentlicht: (2025)
Construction and Evaluation of Mandarin Multimodal Emotional Speech Database
von: Ting, Zhu, et al.
Veröffentlicht: (2024)
von: Ting, Zhu, et al.
Veröffentlicht: (2024)
SAM Audio Judge: A Unified Multimodal Framework for Perceptual Evaluation of Audio Separation
von: Wang, Helin, et al.
Veröffentlicht: (2026)
von: Wang, Helin, et al.
Veröffentlicht: (2026)
AquaSignal: An Integrated Framework for Robust Underwater Acoustic Analysis
von: Panteli, Eirini, et al.
Veröffentlicht: (2025)
von: Panteli, Eirini, et al.
Veröffentlicht: (2025)
AIMDiT: Modality Augmentation and Interaction via Multimodal Dimension Transformation for Emotion Recognition in Conversations
von: Wu, Sheng, et al.
Veröffentlicht: (2024)
von: Wu, Sheng, et al.
Veröffentlicht: (2024)
EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
von: Wu, Zhichao, et al.
Veröffentlicht: (2025)
von: Wu, Zhichao, et al.
Veröffentlicht: (2025)
HuPER: A Human-Inspired Framework for Phonetic Perception
von: Guo, Chenxu, et al.
Veröffentlicht: (2026)
von: Guo, Chenxu, et al.
Veröffentlicht: (2026)
ABHINAYA -- A System for Speech Emotion Recognition In Naturalistic Conditions Challenge
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Characterization of Speech Similarity Between Australian Aboriginal and High-Resource Languages: A Case Study on Dharawal
von: Dang, Ting, et al.
Veröffentlicht: (2025) -
What is Learnt by the LEArnable Front-end (LEAF)? Adapting Per-Channel Energy Normalisation (PCEN) to Noisy Conditions
von: Meng, Hanyu, et al.
Veröffentlicht: (2024) -
Should Audio Front-ends be Adaptive? Comparing Learnable and Adaptive Front-ends
von: Zhang, Qiquan, et al.
Veröffentlicht: (2025) -
Binaural Selective Attention Model for Target Speaker Extraction
von: Meng, Hanyu, et al.
Veröffentlicht: (2024) -
Adaptive Per-Channel Energy Normalization Front-end for Robust Audio Signal Processing
von: Meng, Hanyu, et al.
Veröffentlicht: (2025)