Affectron: Emotional Speech Synthesis with Affective and Contextually Aligned Nonverbal Vocalizations
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Cho, Deok-Hyeon, Oh, Hyung-Seok, Kim, Seung-Bin, Lee, Seong-Whan |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
par: Cho, Deok-Hyeon, et autres
Publié: (2024)
par: Cho, Deok-Hyeon, et autres
Publié: (2024)
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
par: Cho, Deok-Hyeon, et autres
Publié: (2025)
par: Cho, Deok-Hyeon, et autres
Publié: (2025)
EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
par: Cho, Deok-Hyeon, et autres
Publié: (2025)
par: Cho, Deok-Hyeon, et autres
Publié: (2025)
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
par: Cho, Deok-Hyeon, et autres
Publié: (2024)
par: Cho, Deok-Hyeon, et autres
Publié: (2024)
Toward Complex-Valued Neural Networks for Waveform Generation
par: Oh, Hyung-Seok, et autres
Publié: (2026)
par: Oh, Hyung-Seok, et autres
Publié: (2026)
DurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment
par: Oh, Hyung-Seok, et autres
Publié: (2024)
par: Oh, Hyung-Seok, et autres
Publié: (2024)
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis
par: Cha, Jun-Hyeok, et autres
Publié: (2025)
par: Cha, Jun-Hyeok, et autres
Publié: (2025)
Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
par: Kim, Nam-Gyu, et autres
Publié: (2025)
par: Kim, Nam-Gyu, et autres
Publié: (2025)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
par: Oh, Hyung-Seok, et autres
Publié: (2023)
par: Oh, Hyung-Seok, et autres
Publié: (2023)
NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech
par: Borisov, Maksim, et autres
Publié: (2025)
par: Borisov, Maksim, et autres
Publié: (2025)
TranSentence: Speech-to-speech Translation via Language-agnostic Sentence-level Speech Encoding without Language-parallel Data
par: Kim, Seung-Bin, et autres
Publié: (2024)
par: Kim, Seung-Bin, et autres
Publié: (2024)
VibE-SVC: Vibrato Extraction with High-frequency F0 Contour for Singing Voice Conversion
par: Choi, Joon-Seung, et autres
Publié: (2025)
par: Choi, Joon-Seung, et autres
Publié: (2025)
NV-Bench: Benchmark of Nonverbal Vocalization Synthesis for Expressive Text-to-Speech Generation
par: Ni, Qinke, et autres
Publié: (2026)
par: Ni, Qinke, et autres
Publié: (2026)
FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching
par: Yun, Jun-Hak, et autres
Publié: (2025)
par: Yun, Jun-Hak, et autres
Publié: (2025)
Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks
par: Lee, Seo-Hyun, et autres
Publié: (2023)
par: Lee, Seo-Hyun, et autres
Publié: (2023)
MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech
par: Mai, Jialong, et autres
Publié: (2025)
par: Mai, Jialong, et autres
Publié: (2025)
JVNV: A Corpus of Japanese Emotional Speech with Verbal Content and Nonverbal Expressions
par: Xin, Detai, et autres
Publié: (2023)
par: Xin, Detai, et autres
Publié: (2023)
NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations
par: Xue, Liumeng, et autres
Publié: (2026)
par: Xue, Liumeng, et autres
Publié: (2026)
Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
par: Chung, Soo-Whan, et autres
Publié: (2025)
par: Chung, Soo-Whan, et autres
Publié: (2025)
Elderly-Contextual Data Augmentation via Speech Synthesis for Elderly ASR
par: Lee, Minsik, et autres
Publié: (2026)
par: Lee, Minsik, et autres
Publié: (2026)
UniVocal: Unified Speech-Singing Code-Switching Synthesis
par: Shi, Yufei, et autres
Publié: (2026)
par: Shi, Yufei, et autres
Publié: (2026)
LipSody: Lip-to-Speech Synthesis with Enhanced Prosody Consistency
par: Lee, Jaejun, et autres
Publié: (2026)
par: Lee, Jaejun, et autres
Publié: (2026)
Towards Dynamic Neural Communication and Speech Neuroprosthesis Based on Viseme Decoding
par: Park, Ji-Ha, et autres
Publié: (2025)
par: Park, Ji-Ha, et autres
Publié: (2025)
Enhancing Speech Emotion Recognition through Segmental Average Pooling of Self-Supervised Learning Features
par: Hyeon, Jonghwan, et autres
Publié: (2024)
par: Hyeon, Jonghwan, et autres
Publié: (2024)
Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition
par: Zhang, Xu, et autres
Publié: (2026)
par: Zhang, Xu, et autres
Publié: (2026)
AlignCap: Aligning Speech Emotion Captioning to Human Preferences
par: Liang, Ziqi, et autres
Publié: (2024)
par: Liang, Ziqi, et autres
Publié: (2024)
Coding Speech through Vocal Tract Kinematics
par: Cho, Cheol Jun, et autres
Publié: (2024)
par: Cho, Cheol Jun, et autres
Publié: (2024)
Progressive Facial Granularity Aggregation with Bilateral Attribute-based Enhancement for Face-to-Speech Synthesis
par: Jeon, Yejin, et autres
Publié: (2025)
par: Jeon, Yejin, et autres
Publié: (2025)
EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
par: Bian, Weizhen, et autres
Publié: (2024)
par: Bian, Weizhen, et autres
Publié: (2024)
Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
par: Gao, Xiaoxue, et autres
Publié: (2025)
par: Gao, Xiaoxue, et autres
Publié: (2025)
Speaking Without Sound: Multi-speaker Silent Speech Voicing with Facial Inputs Only
par: Lee, Jaejun, et autres
Publié: (2026)
par: Lee, Jaejun, et autres
Publié: (2026)
Adaptive Knowledge Distillation using a Device-Aware Teacher for Low-Complexity Acoustic Scene Classification
par: Jeong, Seung Gyu, et autres
Publié: (2025)
par: Jeong, Seung Gyu, et autres
Publié: (2025)
ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech Synthesis
par: Li, Aoduo, et autres
Publié: (2026)
par: Li, Aoduo, et autres
Publié: (2026)
TASLA: Text-Aligned Speech Tokens with Multiple Layer-Aggregation
par: Hsu, Ming-Hao, et autres
Publié: (2025)
par: Hsu, Ming-Hao, et autres
Publié: (2025)
Hierarchical Control of Emotion Rendering in Speech Synthesis
par: Inoue, Sho, et autres
Publié: (2024)
par: Inoue, Sho, et autres
Publié: (2024)
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
par: Li, Jialu, et autres
Publié: (2024)
par: Li, Jialu, et autres
Publié: (2024)
Audio-Visual Speech Enhancement in Noisy Environments via Emotion-Based Contextual Cues
par: Hussain, Tassadaq, et autres
Publié: (2024)
par: Hussain, Tassadaq, et autres
Publié: (2024)
PiAnnotate: A Web Annotation Tool for Piano Fingering, with a Diagnostic Probe
par: Bae, Joonhyung, et autres
Publié: (2026)
par: Bae, Joonhyung, et autres
Publié: (2026)
Improving Speech-based Emotion Recognition with Contextual Utterance Analysis and LLMs
par: Zhang, Enshi, et autres
Publié: (2024)
par: Zhang, Enshi, et autres
Publié: (2024)
Enhancing Dialogue Speech Recognition with Robust Contextual Awareness via Noise Representation Learning
par: Lee, Wonjun, et autres
Publié: (2024)
par: Lee, Wonjun, et autres
Publié: (2024)
Documents similaires
-
EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
par: Cho, Deok-Hyeon, et autres
Publié: (2024) -
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
par: Cho, Deok-Hyeon, et autres
Publié: (2025) -
EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
par: Cho, Deok-Hyeon, et autres
Publié: (2025) -
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
par: Cho, Deok-Hyeon, et autres
Publié: (2024) -
Toward Complex-Valued Neural Networks for Waveform Generation
par: Oh, Hyung-Seok, et autres
Publié: (2026)