Multi-Microphone Speech Emotion Recognition using the Hierarchical Token-semantic Audio Transformer Architecture
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cohen, Ohad, Hazan, Gershon, Gannot, Sharon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Microphone and Multi-Modal Emotion Recognition in Reverberant Environment
von: Cohen, Ohad, et al.
Veröffentlicht: (2024)
von: Cohen, Ohad, et al.
Veröffentlicht: (2024)
Diffusion-Based Unsupervised Audio-Visual Speech Separation in Noisy Environments with Noise Prior
von: Yemini, Yochai, et al.
Veröffentlicht: (2025)
von: Yemini, Yochai, et al.
Veröffentlicht: (2025)
Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models
von: Zhang, Wenda, et al.
Veröffentlicht: (2026)
von: Zhang, Wenda, et al.
Veröffentlicht: (2026)
Bridging Modalities: Knowledge Distillation and Masked Training for Translating Multi-Modal Emotion Recognition to Uni-Modal, Speech-Only Emotion Recognition
von: Muaz, Muhammad, et al.
Veröffentlicht: (2024)
von: Muaz, Muhammad, et al.
Veröffentlicht: (2024)
Can We Estimate Purchase Intention Based on Zero-shot Speech Emotion Recognition?
von: Nagase, Ryotaro, et al.
Veröffentlicht: (2024)
von: Nagase, Ryotaro, et al.
Veröffentlicht: (2024)
Unsupervised Acoustic Scene Mapping Based on Acoustic Features and Dimensionality Reduction
von: Cohen, Idan, et al.
Veröffentlicht: (2023)
von: Cohen, Idan, et al.
Veröffentlicht: (2023)
SSNAPS: Audio-Visual Separation of Speech and Background Noise with Diffusion Inverse Sampling
von: Yemini, Yochai, et al.
Veröffentlicht: (2026)
von: Yemini, Yochai, et al.
Veröffentlicht: (2026)
DeepEmoNet: Building Machine Learning Models for Automatic Emotion Recognition in Human Speeches
von: Vu, Tai
Veröffentlicht: (2025)
von: Vu, Tai
Veröffentlicht: (2025)
MORE: Multi-Objective Adversarial Attacks on Speech Recognition
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2026)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2026)
Speech Emotion Recognition Using CNN and Its Use Case in Digital Healthcare
von: Nigar, Nishargo
Veröffentlicht: (2024)
von: Nigar, Nishargo
Veröffentlicht: (2024)
Focal Loss based Residual Convolutional Neural Network for Speech Emotion Recognition
von: Tripathi, Suraj, et al.
Veröffentlicht: (2019)
von: Tripathi, Suraj, et al.
Veröffentlicht: (2019)
TRAMBA: A Hybrid Transformer and Mamba Architecture for Practical Audio and Bone Conduction Speech Super Resolution and Enhancement on Mobile and Wearable Platforms
von: Sui, Yueyuan, et al.
Veröffentlicht: (2024)
von: Sui, Yueyuan, et al.
Veröffentlicht: (2024)
Re-ENACT: Reinforcement Learning for Emotional Speech Generation using Actor-Critic Strategy
von: Shankar, Ravi, et al.
Veröffentlicht: (2024)
von: Shankar, Ravi, et al.
Veröffentlicht: (2024)
Enhancing Automatic Speech Recognition Through Integrated Noise Detection Architecture
von: Singh, Karamvir
Veröffentlicht: (2025)
von: Singh, Karamvir
Veröffentlicht: (2025)
Multimodal Audio-based Disease Prediction with Transformer-based Hierarchical Fusion Network
von: Cai, Jinjin, et al.
Veröffentlicht: (2024)
von: Cai, Jinjin, et al.
Veröffentlicht: (2024)
Enhanced Speech Emotion Recognition with Efficient Channel Attention Guided Deep CNN-BiLSTM Framework
von: Kundu, Niloy Kumar, et al.
Veröffentlicht: (2024)
von: Kundu, Niloy Kumar, et al.
Veröffentlicht: (2024)
Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition
von: Wu, Linzhi, et al.
Veröffentlicht: (2026)
von: Wu, Linzhi, et al.
Veröffentlicht: (2026)
Scaling Speech Tokenizers with Diffusion Autoencoders
von: Wang, Yuancheng, et al.
Veröffentlicht: (2026)
von: Wang, Yuancheng, et al.
Veröffentlicht: (2026)
JSQA: Speech Quality Assessment with Perceptually-Inspired Contrastive Pretraining Based on JND Audio Pairs
von: Fan, Junyi, et al.
Veröffentlicht: (2025)
von: Fan, Junyi, et al.
Veröffentlicht: (2025)
LongAudio-RAG: Event-Grounded Question Answering over Multi-Hour Long Audio
von: Vakada, Naveen, et al.
Veröffentlicht: (2026)
von: Vakada, Naveen, et al.
Veröffentlicht: (2026)
Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
Masked Audio Generation using a Single Non-Autoregressive Transformer
von: Ziv, Alon, et al.
Veröffentlicht: (2024)
von: Ziv, Alon, et al.
Veröffentlicht: (2024)
Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization
von: Yang, Yiyuan, et al.
Veröffentlicht: (2025)
von: Yang, Yiyuan, et al.
Veröffentlicht: (2025)
Single-Microphone Speaker Separation and Voice Activity Detection in Noisy and Reverberant Environments
von: Opochinsky, Renana, et al.
Veröffentlicht: (2024)
von: Opochinsky, Renana, et al.
Veröffentlicht: (2024)
RevRIR: Joint Reverberant Speech and Room Impulse Response Embedding using Contrastive Learning with Application to Room Shape Classification
von: Bitterman, Jacob, et al.
Veröffentlicht: (2024)
von: Bitterman, Jacob, et al.
Veröffentlicht: (2024)
A Novel Hybrid Deep Learning Technique for Speech Emotion Detection using Feature Engineering
von: Chowdhury, Shahana Yasmin, et al.
Veröffentlicht: (2025)
von: Chowdhury, Shahana Yasmin, et al.
Veröffentlicht: (2025)
LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
MFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and Hierarchical Cooperative Attention
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024)
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024)
Persian Speech Emotion Recognition by Fine-Tuning Transformers
von: Shayaninasab, Minoo, et al.
Veröffentlicht: (2024)
von: Shayaninasab, Minoo, et al.
Veröffentlicht: (2024)
Speech Enhancement Using Continuous Embeddings of Neural Audio Codec
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
DiffusionRIR: Room Impulse Response Interpolation using Diffusion Models
von: Della Torre, Sagi, et al.
Veröffentlicht: (2025)
von: Della Torre, Sagi, et al.
Veröffentlicht: (2025)
Text-Queried Audio Source Separation via Hierarchical Modeling
von: Yin, Xinlei, et al.
Veröffentlicht: (2025)
von: Yin, Xinlei, et al.
Veröffentlicht: (2025)
Prompt Amplification and Zero-Shot Late Fusion in Audio-Language Models for Speech Emotion Recognition
von: Kataria, Saurabh, et al.
Veröffentlicht: (2026)
von: Kataria, Saurabh, et al.
Veröffentlicht: (2026)
Audio Transformers
von: Verma, Prateek, et al.
Veröffentlicht: (2021)
von: Verma, Prateek, et al.
Veröffentlicht: (2021)
Efficient Parallel Audio Generation using Group Masked Language Modeling
von: Jeong, Myeonghun, et al.
Veröffentlicht: (2024)
von: Jeong, Myeonghun, et al.
Veröffentlicht: (2024)
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
von: Tuncay, Ludovic, et al.
Veröffentlicht: (2025)
von: Tuncay, Ludovic, et al.
Veröffentlicht: (2025)
Efficient Finetuning for Dimensional Speech Emotion Recognition in the Age of Transformers
von: Sampath, Aneesha, et al.
Veröffentlicht: (2025)
von: Sampath, Aneesha, et al.
Veröffentlicht: (2025)
Enhancing Speech Emotion Recognition Through Differentiable Architecture Search
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2023)
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2023)
AMDM-SE: Attention-based Multichannel Diffusion Model for Speech Enhancement
von: Opochinsky, Renana, et al.
Veröffentlicht: (2026)
von: Opochinsky, Renana, et al.
Veröffentlicht: (2026)
A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection
von: Ali, Hashim, et al.
Veröffentlicht: (2026)
von: Ali, Hashim, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Multi-Microphone and Multi-Modal Emotion Recognition in Reverberant Environment
von: Cohen, Ohad, et al.
Veröffentlicht: (2024) -
Diffusion-Based Unsupervised Audio-Visual Speech Separation in Noisy Environments with Noise Prior
von: Yemini, Yochai, et al.
Veröffentlicht: (2025) -
Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models
von: Zhang, Wenda, et al.
Veröffentlicht: (2026) -
Bridging Modalities: Knowledge Distillation and Masked Training for Translating Multi-Modal Emotion Recognition to Uni-Modal, Speech-Only Emotion Recognition
von: Muaz, Muhammad, et al.
Veröffentlicht: (2024) -
Can We Estimate Purchase Intention Based on Zero-shot Speech Emotion Recognition?
von: Nagase, Ryotaro, et al.
Veröffentlicht: (2024)