Re-ENACT: Reinforcement Learning for Emotional Speech Generation using Actor-Critic Strategy
Fuente:
arXiv
Salvato in:
| Autori principali: | Shankar, Ravi, Venkataraman, Archana |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Closer Look at Wav2Vec2 Embeddings for On-Device Single-Channel Speech Enhancement
di: Shankar, Ravi, et al.
Pubblicazione: (2024)
di: Shankar, Ravi, et al.
Pubblicazione: (2024)
Multi-Microphone Speech Emotion Recognition using the Hierarchical Token-semantic Audio Transformer Architecture
di: Cohen, Ohad, et al.
Pubblicazione: (2024)
di: Cohen, Ohad, et al.
Pubblicazione: (2024)
DeepEmoNet: Building Machine Learning Models for Automatic Emotion Recognition in Human Speeches
di: Vu, Tai
Pubblicazione: (2025)
di: Vu, Tai
Pubblicazione: (2025)
A Novel Hybrid Deep Learning Technique for Speech Emotion Detection using Feature Engineering
di: Chowdhury, Shahana Yasmin, et al.
Pubblicazione: (2025)
di: Chowdhury, Shahana Yasmin, et al.
Pubblicazione: (2025)
Can We Estimate Purchase Intention Based on Zero-shot Speech Emotion Recognition?
di: Nagase, Ryotaro, et al.
Pubblicazione: (2024)
di: Nagase, Ryotaro, et al.
Pubblicazione: (2024)
Aligning Generative Speech Enhancement with Perceptual Feedback
di: Li, Haoyang, et al.
Pubblicazione: (2025)
di: Li, Haoyang, et al.
Pubblicazione: (2025)
Improving Generalization of Speech Separation in Real-World Scenarios: Strategies in Simulation, Optimization, and Evaluation
di: Chen, Ke, et al.
Pubblicazione: (2024)
di: Chen, Ke, et al.
Pubblicazione: (2024)
Speech Emotion Recognition Using CNN and Its Use Case in Digital Healthcare
di: Nigar, Nishargo
Pubblicazione: (2024)
di: Nigar, Nishargo
Pubblicazione: (2024)
Focal Loss based Residual Convolutional Neural Network for Speech Emotion Recognition
di: Tripathi, Suraj, et al.
Pubblicazione: (2019)
di: Tripathi, Suraj, et al.
Pubblicazione: (2019)
Bridging Modalities: Knowledge Distillation and Masked Training for Translating Multi-Modal Emotion Recognition to Uni-Modal, Speech-Only Emotion Recognition
di: Muaz, Muhammad, et al.
Pubblicazione: (2024)
di: Muaz, Muhammad, et al.
Pubblicazione: (2024)
Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models
di: Zhang, Wenda, et al.
Pubblicazione: (2026)
di: Zhang, Wenda, et al.
Pubblicazione: (2026)
Enhanced Speech Emotion Recognition with Efficient Channel Attention Guided Deep CNN-BiLSTM Framework
di: Kundu, Niloy Kumar, et al.
Pubblicazione: (2024)
di: Kundu, Niloy Kumar, et al.
Pubblicazione: (2024)
ENACT-Heart -- ENsemble-based Assessment Using CNN and Transformer on Heart Sounds
di: Han, Jiho, et al.
Pubblicazione: (2025)
di: Han, Jiho, et al.
Pubblicazione: (2025)
VINP: Variational Bayesian Inference with Neural Speech Prior for Joint ASR-Effective Speech Dereverberation and Blind RIR Identification
di: Wang, Pengyu, et al.
Pubblicazione: (2025)
di: Wang, Pengyu, et al.
Pubblicazione: (2025)
Single and Few-step Diffusion for Generative Speech Enhancement
di: Lay, Bunlong, et al.
Pubblicazione: (2023)
di: Lay, Bunlong, et al.
Pubblicazione: (2023)
Multi-Metric Preference Alignment for Generative Speech Restoration
di: Zhang, Junan, et al.
Pubblicazione: (2025)
di: Zhang, Junan, et al.
Pubblicazione: (2025)
MORE: Multi-Objective Adversarial Attacks on Speech Recognition
di: Gao, Xiaoxue, et al.
Pubblicazione: (2026)
di: Gao, Xiaoxue, et al.
Pubblicazione: (2026)
Efficient Parallel Audio Generation using Group Masked Language Modeling
di: Jeong, Myeonghun, et al.
Pubblicazione: (2024)
di: Jeong, Myeonghun, et al.
Pubblicazione: (2024)
From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology
di: Li, Haoyang, et al.
Pubblicazione: (2024)
di: Li, Haoyang, et al.
Pubblicazione: (2024)
Something from Nothing: Data Augmentation for Robust Severity Level Estimation of Dysarthric Speech
di: Bae, Jaesung, et al.
Pubblicazione: (2026)
di: Bae, Jaesung, et al.
Pubblicazione: (2026)
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
di: Wang, Yuancheng, et al.
Pubblicazione: (2024)
di: Wang, Yuancheng, et al.
Pubblicazione: (2024)
Self-Supervised Disentangled Representation Learning for Robust Target Speech Extraction
di: Mu, Zhaoxi, et al.
Pubblicazione: (2023)
di: Mu, Zhaoxi, et al.
Pubblicazione: (2023)
JSQA: Speech Quality Assessment with Perceptually-Inspired Contrastive Pretraining Based on JND Audio Pairs
di: Fan, Junyi, et al.
Pubblicazione: (2025)
di: Fan, Junyi, et al.
Pubblicazione: (2025)
Speech Unlearning
di: Cheng, Jiali, et al.
Pubblicazione: (2025)
di: Cheng, Jiali, et al.
Pubblicazione: (2025)
Joint Learning using Mixture-of-Expert-Based Representation for Speech Enhancement and Robust Emotion Recognition
di: Tzeng, Jing-Tong, et al.
Pubblicazione: (2025)
di: Tzeng, Jing-Tong, et al.
Pubblicazione: (2025)
Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech
di: Yang, Dong, et al.
Pubblicazione: (2026)
di: Yang, Dong, et al.
Pubblicazione: (2026)
Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance
di: Hussain, Shehzeen, et al.
Pubblicazione: (2025)
di: Hussain, Shehzeen, et al.
Pubblicazione: (2025)
Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2024)
di: Ulgen, Ismail Rasim, et al.
Pubblicazione: (2024)
JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
di: Ioannides, Georgios, et al.
Pubblicazione: (2025)
di: Ioannides, Georgios, et al.
Pubblicazione: (2025)
Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
di: Fu, Szu-Wei, et al.
Pubblicazione: (2024)
di: Fu, Szu-Wei, et al.
Pubblicazione: (2024)
AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement
di: Zhang, Junan, et al.
Pubblicazione: (2025)
di: Zhang, Junan, et al.
Pubblicazione: (2025)
Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures
di: Ioannides, Georgios, et al.
Pubblicazione: (2026)
di: Ioannides, Georgios, et al.
Pubblicazione: (2026)
TICL+: A Case Study On Speech In-Context Learning for Children's Speech Recognition
di: Zheng, Haolong, et al.
Pubblicazione: (2025)
di: Zheng, Haolong, et al.
Pubblicazione: (2025)
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
di: Wang, Yuancheng, et al.
Pubblicazione: (2025)
di: Wang, Yuancheng, et al.
Pubblicazione: (2025)
SEGAA: A Unified Approach to Predicting Age, Gender, and Emotion in Speech
di: R, Aron, et al.
Pubblicazione: (2024)
di: R, Aron, et al.
Pubblicazione: (2024)
Collaborative Watermarking for Adversarial Speech Synthesis
di: Juvela, Lauri, et al.
Pubblicazione: (2023)
di: Juvela, Lauri, et al.
Pubblicazione: (2023)
Scaling Speech Tokenizers with Diffusion Autoencoders
di: Wang, Yuancheng, et al.
Pubblicazione: (2026)
di: Wang, Yuancheng, et al.
Pubblicazione: (2026)
Domain Adapting Deep Reinforcement Learning for Real-world Speech Emotion Recognition
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2022)
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2022)
Mitigating Unauthorized Speech Synthesis for Voice Protection
di: Zhang, Zhisheng, et al.
Pubblicazione: (2024)
di: Zhang, Zhisheng, et al.
Pubblicazione: (2024)
Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
di: Yang, Chao-Han Huck, et al.
Pubblicazione: (2024)
di: Yang, Chao-Han Huck, et al.
Pubblicazione: (2024)
Documenti analoghi
-
A Closer Look at Wav2Vec2 Embeddings for On-Device Single-Channel Speech Enhancement
di: Shankar, Ravi, et al.
Pubblicazione: (2024) -
Multi-Microphone Speech Emotion Recognition using the Hierarchical Token-semantic Audio Transformer Architecture
di: Cohen, Ohad, et al.
Pubblicazione: (2024) -
DeepEmoNet: Building Machine Learning Models for Automatic Emotion Recognition in Human Speeches
di: Vu, Tai
Pubblicazione: (2025) -
A Novel Hybrid Deep Learning Technique for Speech Emotion Detection using Feature Engineering
di: Chowdhury, Shahana Yasmin, et al.
Pubblicazione: (2025) -
Can We Estimate Purchase Intention Based on Zero-shot Speech Emotion Recognition?
di: Nagase, Ryotaro, et al.
Pubblicazione: (2024)