Progressive Facial Granularity Aggregation with Bilateral Attribute-based Enhancement for Face-to-Speech Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | Jeon, Yejin, Kim, Youngjae, Lee, Jihyun, Kim, Hyounghun, Lee, Gary Geunbae |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech
di: Kim, Youngjae, et al.
Pubblicazione: (2024)
di: Kim, Youngjae, et al.
Pubblicazione: (2024)
Facilitating Personalized TTS for Dysarthric Speakers Using Knowledge Anchoring and Curriculum Learning
di: Jeon, Yejin, et al.
Pubblicazione: (2025)
di: Jeon, Yejin, et al.
Pubblicazione: (2025)
Multi-Level Attention Aggregation for Language-Agnostic Speaker Replication
di: Jeon, Yejin, et al.
Pubblicazione: (2024)
di: Jeon, Yejin, et al.
Pubblicazione: (2024)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
di: Jeon, Yejin, et al.
Pubblicazione: (2024)
di: Jeon, Yejin, et al.
Pubblicazione: (2024)
Leveraging the Interplay Between Syntactic and Acoustic Cues for Optimizing Korean TTS Pause Formation
di: Jeon, Yejin, et al.
Pubblicazione: (2024)
di: Jeon, Yejin, et al.
Pubblicazione: (2024)
An Investigation Into Explainable Audio Hate Speech Detection
di: An, Jinmyeong, et al.
Pubblicazione: (2024)
di: An, Jinmyeong, et al.
Pubblicazione: (2024)
Enhancing Dialogue Speech Recognition with Robust Contextual Awareness via Noise Representation Learning
di: Lee, Wonjun, et al.
Pubblicazione: (2024)
di: Lee, Wonjun, et al.
Pubblicazione: (2024)
PSY-STEP: Structuring Therapeutic Targets and Action Sequences for Proactive Counseling Dialogue Systems
di: Lee, Jihyun, et al.
Pubblicazione: (2026)
di: Lee, Jihyun, et al.
Pubblicazione: (2026)
PanicToCalm: A Proactive Counseling Agent for Panic Attacks
di: Lee, Jihyun, et al.
Pubblicazione: (2025)
di: Lee, Jihyun, et al.
Pubblicazione: (2025)
Mixture-of-Experts with Intermediate CTC Supervision for Accented Speech Recognition
di: Lee, Wonjun, et al.
Pubblicazione: (2026)
di: Lee, Wonjun, et al.
Pubblicazione: (2026)
DyPCL: Dynamic Phoneme-level Contrastive Learning for Dysarthric Speech Recognition
di: Lee, Wonjun, et al.
Pubblicazione: (2025)
di: Lee, Wonjun, et al.
Pubblicazione: (2025)
MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistance
di: Kim, Subin, et al.
Pubblicazione: (2025)
di: Kim, Subin, et al.
Pubblicazione: (2025)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
di: Han, Seungu, et al.
Pubblicazione: (2026)
di: Han, Seungu, et al.
Pubblicazione: (2026)
PicPersona-TOD : A Dataset for Personalizing Utterance Style in Task-Oriented Dialogue with Image Persona
di: Lee, Jihyun, et al.
Pubblicazione: (2025)
di: Lee, Jihyun, et al.
Pubblicazione: (2025)
Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items?
di: Hwang, Seonjeong, et al.
Pubblicazione: (2025)
di: Hwang, Seonjeong, et al.
Pubblicazione: (2025)
Learning When to Translate for Multilingual Reasoning
di: Kang, Deokhyung, et al.
Pubblicazione: (2026)
di: Kang, Deokhyung, et al.
Pubblicazione: (2026)
GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor
di: Lee, Seokgi, et al.
Pubblicazione: (2025)
di: Lee, Seokgi, et al.
Pubblicazione: (2025)
Acoustic Feature Mixup for Balanced Multi-aspect Pronunciation Assessment
di: Do, Heejin, et al.
Pubblicazione: (2024)
di: Do, Heejin, et al.
Pubblicazione: (2024)
Speaking Without Sound: Multi-speaker Silent Speech Voicing with Facial Inputs Only
di: Lee, Jaejun, et al.
Pubblicazione: (2026)
di: Lee, Jaejun, et al.
Pubblicazione: (2026)
Safeguarding RAG Pipelines with GMTP: A Gradient-based Masked Token Probability Method for Poisoned Document Detection
di: Kim, San, et al.
Pubblicazione: (2025)
di: Kim, San, et al.
Pubblicazione: (2025)
Elderly-Contextual Data Augmentation via Speech Synthesis for Elderly ASR
di: Lee, Minsik, et al.
Pubblicazione: (2026)
di: Lee, Minsik, et al.
Pubblicazione: (2026)
ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis
di: Kong, Jungil, et al.
Pubblicazione: (2023)
di: Kong, Jungil, et al.
Pubblicazione: (2023)
Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks
di: Lee, Seo-Hyun, et al.
Pubblicazione: (2023)
di: Lee, Seo-Hyun, et al.
Pubblicazione: (2023)
DroFiT: A Lightweight Band-fused Frequency Attention Toward Real-time UAV Speech Enhancement
di: Lee, Jeongmin, et al.
Pubblicazione: (2025)
di: Lee, Jeongmin, et al.
Pubblicazione: (2025)
Affectron: Emotional Speech Synthesis with Affective and Contextually Aligned Nonverbal Vocalizations
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2026)
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2026)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
di: Jung, Chaeyoung, et al.
Pubblicazione: (2024)
di: Jung, Chaeyoung, et al.
Pubblicazione: (2024)
LipSody: Lip-to-Speech Synthesis with Enhanced Prosody Consistency
di: Lee, Jaejun, et al.
Pubblicazione: (2026)
di: Lee, Jaejun, et al.
Pubblicazione: (2026)
Few-step Adversarial Schrödinger Bridge for Generative Speech Enhancement
di: Han, Seungu, et al.
Pubblicazione: (2025)
di: Han, Seungu, et al.
Pubblicazione: (2025)
TASLA: Text-Aligned Speech Tokens with Multiple Layer-Aggregation
di: Hsu, Ming-Hao, et al.
Pubblicazione: (2025)
di: Hsu, Ming-Hao, et al.
Pubblicazione: (2025)
DISPATCH: Distilling Selective Patches for Speech Enhancement
di: Kim, Dohwan, et al.
Pubblicazione: (2025)
di: Kim, Dohwan, et al.
Pubblicazione: (2025)
SAGE-LD: Towards Scalable and Generalizable End-to-End Language Diarization via Simulated Data Augmentation
di: Lee, Sangmin, et al.
Pubblicazione: (2025)
di: Lee, Sangmin, et al.
Pubblicazione: (2025)
RoVo: Robust Voice Protection Against Unauthorized Speech Synthesis with Embedding-Level Perturbations
di: Kim, Seungmin, et al.
Pubblicazione: (2025)
di: Kim, Seungmin, et al.
Pubblicazione: (2025)
Self-Correcting Code Generation Using Small Language Models
di: Cho, Jeonghun, et al.
Pubblicazione: (2025)
di: Cho, Jeonghun, et al.
Pubblicazione: (2025)
A Multi-Agent Framework for Feature-Constrained Difficulty Control in Reading Comprehension Item Generation
di: Hwang, Seonjeong, et al.
Pubblicazione: (2026)
di: Hwang, Seonjeong, et al.
Pubblicazione: (2026)
Raon-Speech Technical Report
di: Kim, Beomsoo, et al.
Pubblicazione: (2026)
di: Kim, Beomsoo, et al.
Pubblicazione: (2026)
Inference is All You Need: Self Example Retriever for Cross-domain Dialogue State Tracking with ChatGPT
di: Lee, Jihyun, et al.
Pubblicazione: (2024)
di: Lee, Jihyun, et al.
Pubblicazione: (2024)
Diffusion-based Frameworks for Unsupervised Speech Enhancement
di: Ayilo, Jean-Eudes, et al.
Pubblicazione: (2026)
di: Ayilo, Jean-Eudes, et al.
Pubblicazione: (2026)
Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling
di: Zhang, Leying, et al.
Pubblicazione: (2024)
di: Zhang, Leying, et al.
Pubblicazione: (2024)
Modality-Specific Speech Enhancement and Noise-Adaptive Fusion for Acoustic and Body-Conduction Microphone Framework
di: Kim, Yunsik, et al.
Pubblicazione: (2025)
di: Kim, Yunsik, et al.
Pubblicazione: (2025)
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
di: Chao, Rong, et al.
Pubblicazione: (2025)
di: Chao, Rong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech
di: Kim, Youngjae, et al.
Pubblicazione: (2024) -
Facilitating Personalized TTS for Dysarthric Speakers Using Knowledge Anchoring and Curriculum Learning
di: Jeon, Yejin, et al.
Pubblicazione: (2025) -
Multi-Level Attention Aggregation for Language-Agnostic Speaker Replication
di: Jeon, Yejin, et al.
Pubblicazione: (2024) -
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
di: Jeon, Yejin, et al.
Pubblicazione: (2024) -
Leveraging the Interplay Between Syntactic and Acoustic Cues for Optimizing Korean TTS Pause Formation
di: Jeon, Yejin, et al.
Pubblicazione: (2024)