Progressive Facial Granularity Aggregation with Bilateral Attribute-based Enhancement for Face-to-Speech Synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jeon, Yejin, Kim, Youngjae, Lee, Jihyun, Kim, Hyounghun, Lee, Gary Geunbae |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech
von: Kim, Youngjae, et al.
Veröffentlicht: (2024)
von: Kim, Youngjae, et al.
Veröffentlicht: (2024)
Facilitating Personalized TTS for Dysarthric Speakers Using Knowledge Anchoring and Curriculum Learning
von: Jeon, Yejin, et al.
Veröffentlicht: (2025)
von: Jeon, Yejin, et al.
Veröffentlicht: (2025)
Multi-Level Attention Aggregation for Language-Agnostic Speaker Replication
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
Leveraging the Interplay Between Syntactic and Acoustic Cues for Optimizing Korean TTS Pause Formation
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
An Investigation Into Explainable Audio Hate Speech Detection
von: An, Jinmyeong, et al.
Veröffentlicht: (2024)
von: An, Jinmyeong, et al.
Veröffentlicht: (2024)
Enhancing Dialogue Speech Recognition with Robust Contextual Awareness via Noise Representation Learning
von: Lee, Wonjun, et al.
Veröffentlicht: (2024)
von: Lee, Wonjun, et al.
Veröffentlicht: (2024)
PSY-STEP: Structuring Therapeutic Targets and Action Sequences for Proactive Counseling Dialogue Systems
von: Lee, Jihyun, et al.
Veröffentlicht: (2026)
von: Lee, Jihyun, et al.
Veröffentlicht: (2026)
PanicToCalm: A Proactive Counseling Agent for Panic Attacks
von: Lee, Jihyun, et al.
Veröffentlicht: (2025)
von: Lee, Jihyun, et al.
Veröffentlicht: (2025)
Mixture-of-Experts with Intermediate CTC Supervision for Accented Speech Recognition
von: Lee, Wonjun, et al.
Veröffentlicht: (2026)
von: Lee, Wonjun, et al.
Veröffentlicht: (2026)
DyPCL: Dynamic Phoneme-level Contrastive Learning for Dysarthric Speech Recognition
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistance
von: Kim, Subin, et al.
Veröffentlicht: (2025)
von: Kim, Subin, et al.
Veröffentlicht: (2025)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
von: Han, Seungu, et al.
Veröffentlicht: (2026)
von: Han, Seungu, et al.
Veröffentlicht: (2026)
PicPersona-TOD : A Dataset for Personalizing Utterance Style in Task-Oriented Dialogue with Image Persona
von: Lee, Jihyun, et al.
Veröffentlicht: (2025)
von: Lee, Jihyun, et al.
Veröffentlicht: (2025)
Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items?
von: Hwang, Seonjeong, et al.
Veröffentlicht: (2025)
von: Hwang, Seonjeong, et al.
Veröffentlicht: (2025)
Learning When to Translate for Multilingual Reasoning
von: Kang, Deokhyung, et al.
Veröffentlicht: (2026)
von: Kang, Deokhyung, et al.
Veröffentlicht: (2026)
GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor
von: Lee, Seokgi, et al.
Veröffentlicht: (2025)
von: Lee, Seokgi, et al.
Veröffentlicht: (2025)
Acoustic Feature Mixup for Balanced Multi-aspect Pronunciation Assessment
von: Do, Heejin, et al.
Veröffentlicht: (2024)
von: Do, Heejin, et al.
Veröffentlicht: (2024)
Speaking Without Sound: Multi-speaker Silent Speech Voicing with Facial Inputs Only
von: Lee, Jaejun, et al.
Veröffentlicht: (2026)
von: Lee, Jaejun, et al.
Veröffentlicht: (2026)
Safeguarding RAG Pipelines with GMTP: A Gradient-based Masked Token Probability Method for Poisoned Document Detection
von: Kim, San, et al.
Veröffentlicht: (2025)
von: Kim, San, et al.
Veröffentlicht: (2025)
Elderly-Contextual Data Augmentation via Speech Synthesis for Elderly ASR
von: Lee, Minsik, et al.
Veröffentlicht: (2026)
von: Lee, Minsik, et al.
Veröffentlicht: (2026)
ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis
von: Kong, Jungil, et al.
Veröffentlicht: (2023)
von: Kong, Jungil, et al.
Veröffentlicht: (2023)
Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks
von: Lee, Seo-Hyun, et al.
Veröffentlicht: (2023)
von: Lee, Seo-Hyun, et al.
Veröffentlicht: (2023)
DroFiT: A Lightweight Band-fused Frequency Attention Toward Real-time UAV Speech Enhancement
von: Lee, Jeongmin, et al.
Veröffentlicht: (2025)
von: Lee, Jeongmin, et al.
Veröffentlicht: (2025)
Affectron: Emotional Speech Synthesis with Affective and Contextually Aligned Nonverbal Vocalizations
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2026)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2026)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
LipSody: Lip-to-Speech Synthesis with Enhanced Prosody Consistency
von: Lee, Jaejun, et al.
Veröffentlicht: (2026)
von: Lee, Jaejun, et al.
Veröffentlicht: (2026)
Few-step Adversarial Schrödinger Bridge for Generative Speech Enhancement
von: Han, Seungu, et al.
Veröffentlicht: (2025)
von: Han, Seungu, et al.
Veröffentlicht: (2025)
TASLA: Text-Aligned Speech Tokens with Multiple Layer-Aggregation
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2025)
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2025)
DISPATCH: Distilling Selective Patches for Speech Enhancement
von: Kim, Dohwan, et al.
Veröffentlicht: (2025)
von: Kim, Dohwan, et al.
Veröffentlicht: (2025)
SAGE-LD: Towards Scalable and Generalizable End-to-End Language Diarization via Simulated Data Augmentation
von: Lee, Sangmin, et al.
Veröffentlicht: (2025)
von: Lee, Sangmin, et al.
Veröffentlicht: (2025)
RoVo: Robust Voice Protection Against Unauthorized Speech Synthesis with Embedding-Level Perturbations
von: Kim, Seungmin, et al.
Veröffentlicht: (2025)
von: Kim, Seungmin, et al.
Veröffentlicht: (2025)
Self-Correcting Code Generation Using Small Language Models
von: Cho, Jeonghun, et al.
Veröffentlicht: (2025)
von: Cho, Jeonghun, et al.
Veröffentlicht: (2025)
A Multi-Agent Framework for Feature-Constrained Difficulty Control in Reading Comprehension Item Generation
von: Hwang, Seonjeong, et al.
Veröffentlicht: (2026)
von: Hwang, Seonjeong, et al.
Veröffentlicht: (2026)
Raon-Speech Technical Report
von: Kim, Beomsoo, et al.
Veröffentlicht: (2026)
von: Kim, Beomsoo, et al.
Veröffentlicht: (2026)
Inference is All You Need: Self Example Retriever for Cross-domain Dialogue State Tracking with ChatGPT
von: Lee, Jihyun, et al.
Veröffentlicht: (2024)
von: Lee, Jihyun, et al.
Veröffentlicht: (2024)
Diffusion-based Frameworks for Unsupervised Speech Enhancement
von: Ayilo, Jean-Eudes, et al.
Veröffentlicht: (2026)
von: Ayilo, Jean-Eudes, et al.
Veröffentlicht: (2026)
Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling
von: Zhang, Leying, et al.
Veröffentlicht: (2024)
von: Zhang, Leying, et al.
Veröffentlicht: (2024)
Modality-Specific Speech Enhancement and Noise-Adaptive Fusion for Acoustic and Body-Conduction Microphone Framework
von: Kim, Yunsik, et al.
Veröffentlicht: (2025)
von: Kim, Yunsik, et al.
Veröffentlicht: (2025)
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
von: Chao, Rong, et al.
Veröffentlicht: (2025)
von: Chao, Rong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech
von: Kim, Youngjae, et al.
Veröffentlicht: (2024) -
Facilitating Personalized TTS for Dysarthric Speakers Using Knowledge Anchoring and Curriculum Learning
von: Jeon, Yejin, et al.
Veröffentlicht: (2025) -
Multi-Level Attention Aggregation for Language-Agnostic Speaker Replication
von: Jeon, Yejin, et al.
Veröffentlicht: (2024) -
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
von: Jeon, Yejin, et al.
Veröffentlicht: (2024) -
Leveraging the Interplay Between Syntactic and Acoustic Cues for Optimizing Korean TTS Pause Formation
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)