A conversational gesture synthesis system based on emotions and semantics
Fuente:
arXiv
Guardado en:
| Autor principal: | Hoang-Minh, Thanh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Head, posture, and full-body gestures in unscripted dyadic conversations in noise
por: Hládek, Ľuboš, et al.
Publicado: (2025)
por: Hládek, Ľuboš, et al.
Publicado: (2025)
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
por: Torgashov, Nikita, et al.
Publicado: (2025)
por: Torgashov, Nikita, et al.
Publicado: (2025)
LLAMAPIE: Proactive In-Ear Conversation Assistants
por: Chen, Tuochao, et al.
Publicado: (2025)
por: Chen, Tuochao, et al.
Publicado: (2025)
Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education
por: Snoubara, Abdul Aziz, et al.
Publicado: (2026)
por: Snoubara, Abdul Aziz, et al.
Publicado: (2026)
Spontaneous Informal Speech Dataset for Punctuation Restoration
por: Liu, Xing Yi, et al.
Publicado: (2024)
por: Liu, Xing Yi, et al.
Publicado: (2024)
AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
por: Kawamura, Kazuki, et al.
Publicado: (2024)
por: Kawamura, Kazuki, et al.
Publicado: (2024)
Loop Copilot: Conducting AI Ensembles for Music Generation and Iterative Editing
por: Zhang, Yixiao, et al.
Publicado: (2023)
por: Zhang, Yixiao, et al.
Publicado: (2023)
VoXtream2: Full-stream TTS with dynamic speaking rate control
por: Torgashov, Nikita, et al.
Publicado: (2026)
por: Torgashov, Nikita, et al.
Publicado: (2026)
Literary and Colloquial Tamil Dialect Identification
por: Nanmalar, M., et al.
Publicado: (2024)
por: Nanmalar, M., et al.
Publicado: (2024)
Efficient Ensemble for Multimodal Punctuation Restoration using Time-Delay Neural Network
por: Liu, Xing Yi, et al.
Publicado: (2023)
por: Liu, Xing Yi, et al.
Publicado: (2023)
Quality Audio Prototyping: a prototype system for unified sound retrieval and procedural generation
por: Garcia, Nelly, et al.
Publicado: (2026)
por: Garcia, Nelly, et al.
Publicado: (2026)
DOO-RE: A dataset of ambient sensors in a meeting room for activity recognition
por: Kim, Hyunju, et al.
Publicado: (2024)
por: Kim, Hyunju, et al.
Publicado: (2024)
Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach
por: Sankey-Olsen, Cuno, et al.
Publicado: (2025)
por: Sankey-Olsen, Cuno, et al.
Publicado: (2025)
Clip-TTS: Contrastive Text-content and Mel-spectrogram, A High-Quality Text-to-Speech Method based on Contextual Semantic Understanding
por: Liu, Tianyun
Publicado: (2025)
por: Liu, Tianyun
Publicado: (2025)
Open vocabulary keyword spotting through transfer learning from speech synthesis
por: V, Kesavaraj, et al.
Publicado: (2024)
por: V, Kesavaraj, et al.
Publicado: (2024)
SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures
por: Yuan, Kuang, et al.
Publicado: (2025)
por: Yuan, Kuang, et al.
Publicado: (2025)
Improving AI-generated music with user-guided training
por: Singh, Vishwa Mohan, et al.
Publicado: (2025)
por: Singh, Vishwa Mohan, et al.
Publicado: (2025)
Human Feedback Driven Dynamic Speech Emotion Recognition
por: Fedorov, Ilya, et al.
Publicado: (2025)
por: Fedorov, Ilya, et al.
Publicado: (2025)
Optimizing Multilingual Text-To-Speech with Accents & Emotions
por: Pawar, Pranav, et al.
Publicado: (2025)
por: Pawar, Pranav, et al.
Publicado: (2025)
Artificial Neural Networks to Recognize Speakers Division from Continuous Bengali Speech
por: Ali, Hasmot, et al.
Publicado: (2024)
por: Ali, Hasmot, et al.
Publicado: (2024)
Harnessing Smartwatch Microphone Sensors for Cough Detection and Classification
por: Jaiswal, Pranay, et al.
Publicado: (2024)
por: Jaiswal, Pranay, et al.
Publicado: (2024)
Voice Passing : a Non-Binary Voice Gender Prediction System for evaluating Transgender voice transition
por: Doukhan, David, et al.
Publicado: (2024)
por: Doukhan, David, et al.
Publicado: (2024)
Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese
por: Wang, Xihuai, et al.
Publicado: (2025)
por: Wang, Xihuai, et al.
Publicado: (2025)
SCDiar: a streaming diarization system based on speaker change detection and speech recognition
por: Zheng, Naijun, et al.
Publicado: (2025)
por: Zheng, Naijun, et al.
Publicado: (2025)
Call2Instruct: Automated Pipeline for Generating Q&A Datasets from Call Center Recordings for LLM Fine-Tuning
por: Echeverria, Alex, et al.
Publicado: (2025)
por: Echeverria, Alex, et al.
Publicado: (2025)
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
por: Cheng, Xize, et al.
Publicado: (2025)
por: Cheng, Xize, et al.
Publicado: (2025)
Are Expressions for Music Emotions the Same Across Cultures?
por: Celen, Elif, et al.
Publicado: (2025)
por: Celen, Elif, et al.
Publicado: (2025)
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
por: Wang, Dingdong, et al.
Publicado: (2025)
por: Wang, Dingdong, et al.
Publicado: (2025)
Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection
por: Inoue, Koji, et al.
Publicado: (2024)
por: Inoue, Koji, et al.
Publicado: (2024)
Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection
por: Inoue, Koji, et al.
Publicado: (2024)
por: Inoue, Koji, et al.
Publicado: (2024)
Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction
por: Jeon, Hyunbae, et al.
Publicado: (2024)
por: Jeon, Hyunbae, et al.
Publicado: (2024)
The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era
por: Zhao, Zhixian, et al.
Publicado: (2026)
por: Zhao, Zhixian, et al.
Publicado: (2026)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
por: Sharma, Roshan, et al.
Publicado: (2024)
por: Sharma, Roshan, et al.
Publicado: (2024)
Enhancing AAC Software for Dysarthric Speakers in e-Health Settings: An Evaluation Using TORGO
por: Hui, Macarious, et al.
Publicado: (2024)
por: Hui, Macarious, et al.
Publicado: (2024)
MCMChaos: Improvising Rap Music with MCMC Methods and Chaos Theory
por: Kimelman, Robert G.
Publicado: (2024)
por: Kimelman, Robert G.
Publicado: (2024)
Detecting the terminality of speech-turn boundary for spoken interactions in French TV and Radio content
por: Uro, Rémi, et al.
Publicado: (2024)
por: Uro, Rémi, et al.
Publicado: (2024)
Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model
por: Fu, Yu-Kuan, et al.
Publicado: (2024)
por: Fu, Yu-Kuan, et al.
Publicado: (2024)
EmoHeal: An End-to-End System for Personalized Therapeutic Music Retrieval from Fine-grained Emotions
por: Wan, Xinchen, et al.
Publicado: (2025)
por: Wan, Xinchen, et al.
Publicado: (2025)
Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
por: Xie, Zhifei, et al.
Publicado: (2024)
por: Xie, Zhifei, et al.
Publicado: (2024)
The language of sound search: Examining User Queries in Audio Search Engines
por: Weck, Benno, et al.
Publicado: (2024)
por: Weck, Benno, et al.
Publicado: (2024)
Ejemplares similares
-
Head, posture, and full-body gestures in unscripted dyadic conversations in noise
por: Hládek, Ľuboš, et al.
Publicado: (2025) -
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
por: Torgashov, Nikita, et al.
Publicado: (2025) -
LLAMAPIE: Proactive In-Ear Conversation Assistants
por: Chen, Tuochao, et al.
Publicado: (2025) -
Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education
por: Snoubara, Abdul Aziz, et al.
Publicado: (2026) -
Spontaneous Informal Speech Dataset for Punctuation Restoration
por: Liu, Xing Yi, et al.
Publicado: (2024)