JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis
Fuente:
arXiv
Guardado en:
| Autores principales: | Cha, Jun-Hyeok, Kim, Seung-Bin, Oh, Hyung-Seok, Lee, Seong-Whan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
por: Cho, Deok-Hyeon, et al.
Publicado: (2025)
por: Cho, Deok-Hyeon, et al.
Publicado: (2025)
EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
por: Cho, Deok-Hyeon, et al.
Publicado: (2024)
por: Cho, Deok-Hyeon, et al.
Publicado: (2024)
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
por: Cho, Deok-Hyeon, et al.
Publicado: (2025)
por: Cho, Deok-Hyeon, et al.
Publicado: (2025)
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
por: Cho, Deok-Hyeon, et al.
Publicado: (2024)
por: Cho, Deok-Hyeon, et al.
Publicado: (2024)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
por: Oh, Hyung-Seok, et al.
Publicado: (2023)
por: Oh, Hyung-Seok, et al.
Publicado: (2023)
DurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment
por: Oh, Hyung-Seok, et al.
Publicado: (2024)
por: Oh, Hyung-Seok, et al.
Publicado: (2024)
VibE-SVC: Vibrato Extraction with High-frequency F0 Contour for Singing Voice Conversion
por: Choi, Joon-Seung, et al.
Publicado: (2025)
por: Choi, Joon-Seung, et al.
Publicado: (2025)
TranSentence: Speech-to-speech Translation via Language-agnostic Sentence-level Speech Encoding without Language-parallel Data
por: Kim, Seung-Bin, et al.
Publicado: (2024)
por: Kim, Seung-Bin, et al.
Publicado: (2024)
Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
por: Chung, Soo-Whan, et al.
Publicado: (2025)
por: Chung, Soo-Whan, et al.
Publicado: (2025)
Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
por: Kim, Nam-Gyu, et al.
Publicado: (2025)
por: Kim, Nam-Gyu, et al.
Publicado: (2025)
Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks
por: Lee, Seo-Hyun, et al.
Publicado: (2023)
por: Lee, Seo-Hyun, et al.
Publicado: (2023)
FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching
por: Yun, Jun-Hak, et al.
Publicado: (2025)
por: Yun, Jun-Hak, et al.
Publicado: (2025)
TF-CorrNet: Leveraging Spatial Correlation for Continuous Speech Separation
por: Shin, Ui-Hyeop, et al.
Publicado: (2025)
por: Shin, Ui-Hyeop, et al.
Publicado: (2025)
EMO-SUPERB: An In-depth Look at Speech Emotion Recognition
por: Wu, Haibin, et al.
Publicado: (2024)
por: Wu, Haibin, et al.
Publicado: (2024)
Seeing the Context: Rich Visual Context-Aware Speech Recognition via Multimodal Reasoning
por: Tian, Wenjie, et al.
Publicado: (2026)
por: Tian, Wenjie, et al.
Publicado: (2026)
Speech Emotion Recognition with ASR Integration
por: Li, Yuanchao
Publicado: (2026)
por: Li, Yuanchao
Publicado: (2026)
Dataset-Distillation Generative Model for Speech Emotion Recognition
por: Ritter-Gutierrez, Fabian, et al.
Publicado: (2024)
por: Ritter-Gutierrez, Fabian, et al.
Publicado: (2024)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
por: Shen, Siyuan, et al.
Publicado: (2024)
por: Shen, Siyuan, et al.
Publicado: (2024)
THAI Speech Emotion Recognition (THAI-SER) corpus
por: Wongpithayadisai, Jilamika, et al.
Publicado: (2025)
por: Wongpithayadisai, Jilamika, et al.
Publicado: (2025)
Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
por: Sun, Haoqin, et al.
Publicado: (2024)
por: Sun, Haoqin, et al.
Publicado: (2024)
RSET: Remapping-based Sorting Method for Emotion Transfer Speech Synthesis
por: Shi, Haoxiang, et al.
Publicado: (2024)
por: Shi, Haoxiang, et al.
Publicado: (2024)
CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation
por: Oh, Hyunwoo, et al.
Publicado: (2025)
por: Oh, Hyunwoo, et al.
Publicado: (2025)
Statistical Beamformer Exploiting Non-stationarity and Sparsity with Spatially Constrained ICA for Robust Speech Recognition
por: Shin, Ui-Hyeop, et al.
Publicado: (2023)
por: Shin, Ui-Hyeop, et al.
Publicado: (2023)
Hierarchical Control of Emotion Rendering in Speech Synthesis
por: Inoue, Sho, et al.
Publicado: (2024)
por: Inoue, Sho, et al.
Publicado: (2024)
LLM supervised Pre-training for Multimodal Emotion Recognition in Conversations
por: Dutta, Soumya, et al.
Publicado: (2025)
por: Dutta, Soumya, et al.
Publicado: (2025)
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
por: Yen, Hao, et al.
Publicado: (2024)
por: Yen, Hao, et al.
Publicado: (2024)
Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs
por: Zhang, Wenyu, et al.
Publicado: (2025)
por: Zhang, Wenyu, et al.
Publicado: (2025)
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
por: Bai, Ye, et al.
Publicado: (2024)
por: Bai, Ye, et al.
Publicado: (2024)
PCQ: Emotion Recognition in Speech via Progressive Channel Querying
por: Wang, Xincheng, et al.
Publicado: (2024)
por: Wang, Xincheng, et al.
Publicado: (2024)
Testing Correctness, Fairness, and Robustness of Speech Emotion Recognition Models
por: Derington, Anna, et al.
Publicado: (2023)
por: Derington, Anna, et al.
Publicado: (2023)
SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios
por: Bukhari, Hazim, et al.
Publicado: (2024)
por: Bukhari, Hazim, et al.
Publicado: (2024)
Long-Context Speech Synthesis with Context-Aware Memory
por: Li, Zhipeng, et al.
Publicado: (2025)
por: Li, Zhipeng, et al.
Publicado: (2025)
Unifying Speech Recognition, Synthesis and Conversion with Autoregressive Transformers
por: Cai, Runyuan, et al.
Publicado: (2026)
por: Cai, Runyuan, et al.
Publicado: (2026)
Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
por: Zhao, Yan, et al.
Publicado: (2024)
por: Zhao, Yan, et al.
Publicado: (2024)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
por: Inoue, Sho, et al.
Publicado: (2024)
por: Inoue, Sho, et al.
Publicado: (2024)
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
por: Gao, Xiaoxue, et al.
Publicado: (2024)
por: Gao, Xiaoxue, et al.
Publicado: (2024)
MSP-Conversation: A Corpus for Naturalistic, Time-Continuous Emotion Recognition
por: Martinez-Lucas, Luz, et al.
Publicado: (2026)
por: Martinez-Lucas, Luz, et al.
Publicado: (2026)
Towards Dynamic Neural Communication and Speech Neuroprosthesis Based on Viseme Decoding
por: Park, Ji-Ha, et al.
Publicado: (2025)
por: Park, Ji-Ha, et al.
Publicado: (2025)
Machine Unlearning in Speech Emotion Recognition via Forget Set Alone
por: Ren, Zhao, et al.
Publicado: (2025)
por: Ren, Zhao, et al.
Publicado: (2025)
Semantic-Emotional Resonance Embedding: A Semi-Supervised Paradigm for Cross-Lingual Speech Emotion Recognition
por: Zhao, Ya, et al.
Publicado: (2026)
por: Zhao, Ya, et al.
Publicado: (2026)
Ejemplares similares
-
EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
por: Cho, Deok-Hyeon, et al.
Publicado: (2025) -
EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
por: Cho, Deok-Hyeon, et al.
Publicado: (2024) -
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
por: Cho, Deok-Hyeon, et al.
Publicado: (2025) -
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
por: Cho, Deok-Hyeon, et al.
Publicado: (2024) -
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
por: Oh, Hyung-Seok, et al.
Publicado: (2023)