Leveraging WaveNet for Dynamic Listening Head Modeling from Speech
Fuente:
arXiv
Salvato in:
| Autori principali: | Nguyen, Minh-Duc, Yang, Hyung-Jeong, Kim, Seung-Won, Shin, Ji-Eun, Kim, Soo-Hyung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Transformer with Leveraged Masked Autoencoder for video-based Pain Assessment
di: Nguyen, Minh-Duc, et al.
Pubblicazione: (2024)
di: Nguyen, Minh-Duc, et al.
Pubblicazione: (2024)
Anatomical Attention Alignment representation for Radiology Report Generation
di: Nguyen, Quang Vinh, et al.
Pubblicazione: (2025)
di: Nguyen, Quang Vinh, et al.
Pubblicazione: (2025)
ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise Diffusion
di: Vo, Hoang-Son, et al.
Pubblicazione: (2025)
di: Vo, Hoang-Son, et al.
Pubblicazione: (2025)
Latent Behavior Diffusion for Sequential Reaction Generation in Dyadic Setting
di: Nguyen, Minh-Duc, et al.
Pubblicazione: (2025)
di: Nguyen, Minh-Duc, et al.
Pubblicazione: (2025)
WaveNets: Wavelet Channel Attention Networks
di: Salman, Hadi, et al.
Pubblicazione: (2022)
di: Salman, Hadi, et al.
Pubblicazione: (2022)
Learning Phonetic Context-Dependent Viseme for Enhancing Speech-Driven 3D Facial Animation
di: Kim, Hyung Kyu, et al.
Pubblicazione: (2025)
di: Kim, Hyung Kyu, et al.
Pubblicazione: (2025)
Conditional Diffusion Model for Longitudinal Medical Image Generation
di: Dao, Duy-Phuong, et al.
Pubblicazione: (2024)
di: Dao, Duy-Phuong, et al.
Pubblicazione: (2024)
WISE-FUSE: Efficient Whole Slide Image Encoding via Coarse-to-Fine Patch Selection with VLM and LLM Knowledge Fusion
di: Shin, Yonghan, et al.
Pubblicazione: (2025)
di: Shin, Yonghan, et al.
Pubblicazione: (2025)
KAN-Based Fusion of Dual-Domain for Audio-Driven Facial Landmarks Generation
di: Vo-Thanh, Hoang-Son, et al.
Pubblicazione: (2024)
di: Vo-Thanh, Hoang-Son, et al.
Pubblicazione: (2024)
Rethinking Top Probability from Multi-view for Distracted Driver Behaviour Localization
di: Nguyen, Quang Vinh, et al.
Pubblicazione: (2024)
di: Nguyen, Quang Vinh, et al.
Pubblicazione: (2024)
MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization
di: Kim, Hyung Kyu, et al.
Pubblicazione: (2025)
di: Kim, Hyung Kyu, et al.
Pubblicazione: (2025)
Temporal In-Context Fine-Tuning with Temporal Reasoning for Versatile Control of Video Diffusion Models
di: Kim, Kinam, et al.
Pubblicazione: (2025)
di: Kim, Kinam, et al.
Pubblicazione: (2025)
Bidirectional Regression for Monocular 6DoF Head Pose Estimation and Reference System Alignment
di: Chun, Sungho, et al.
Pubblicazione: (2024)
di: Chun, Sungho, et al.
Pubblicazione: (2024)
BoIR: Box-Supervised Instance Representation for Multi-Person Pose Estimation
di: Jeong, Uyoung, et al.
Pubblicazione: (2023)
di: Jeong, Uyoung, et al.
Pubblicazione: (2023)
Cross-domain Denoising for Low-dose Multi-frame Spiral Computed Tomography
di: Lu, Yucheng, et al.
Pubblicazione: (2023)
di: Lu, Yucheng, et al.
Pubblicazione: (2023)
ConcreTizer: Model Inversion Attack via Occupancy Classification and Dispersion Control for 3D Point Cloud Restoration
di: Kim, Youngseok, et al.
Pubblicazione: (2025)
di: Kim, Youngseok, et al.
Pubblicazione: (2025)
Skip-WaveNet: A Wavelet based Multi-scale Architecture to Trace Snow Layers in Radar Echograms
di: Varshney, Debvrat, et al.
Pubblicazione: (2023)
di: Varshney, Debvrat, et al.
Pubblicazione: (2023)
MonoWAD: Weather-Adaptive Diffusion Model for Robust Monocular 3D Object Detection
di: Oh, Youngmin, et al.
Pubblicazione: (2024)
di: Oh, Youngmin, et al.
Pubblicazione: (2024)
Leveraging Spatial Attention and Edge Context for Optimized Feature Selection in Visual Localization
di: Istighfarin, Nanda Febri, et al.
Pubblicazione: (2024)
di: Istighfarin, Nanda Febri, et al.
Pubblicazione: (2024)
Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models
di: Jin, Hyundong, et al.
Pubblicazione: (2025)
di: Jin, Hyundong, et al.
Pubblicazione: (2025)
See It All: Contextualized Late Aggregation for 3D Dense Captioning
di: Kim, Minjung, et al.
Pubblicazione: (2024)
di: Kim, Minjung, et al.
Pubblicazione: (2024)
WaveNet-SF: A Hybrid Network for Retinal Disease Detection Based on Wavelet Transform in Spatial-Frequency Domain
di: Cheng, Jilan, et al.
Pubblicazione: (2025)
di: Cheng, Jilan, et al.
Pubblicazione: (2025)
PoseBH: Prototypical Multi-Dataset Training Beyond Human Pose Estimation
di: Jeong, Uyoung, et al.
Pubblicazione: (2025)
di: Jeong, Uyoung, et al.
Pubblicazione: (2025)
TwinLiteNet+: An Enhanced Multi-Task Segmentation Model for Autonomous Driving
di: Che, Quang-Huy, et al.
Pubblicazione: (2024)
di: Che, Quang-Huy, et al.
Pubblicazione: (2024)
Enhanced fringe-to-phase framework using deep learning
di: Kim, Won-Hoe, et al.
Pubblicazione: (2024)
di: Kim, Won-Hoe, et al.
Pubblicazione: (2024)
From Tokens to Photons: Test-Time Physical Prompting for Vision-Language Models
di: Im, Boyeong, et al.
Pubblicazione: (2025)
di: Im, Boyeong, et al.
Pubblicazione: (2025)
Bi-directional Contextual Attention for 3D Dense Captioning
di: Kim, Minjung, et al.
Pubblicazione: (2024)
di: Kim, Minjung, et al.
Pubblicazione: (2024)
THOM: Generating Physically Plausible Hand-Object Meshes From Text
di: Jeong, Uyoung, et al.
Pubblicazione: (2026)
di: Jeong, Uyoung, et al.
Pubblicazione: (2026)
CLIMB: Controllable Longitudinal Brain Image Generation using Mamba-based Latent Diffusion Model and Gaussian-aligned Autoencoder
di: Dao, Duy-Phuong, et al.
Pubblicazione: (2026)
di: Dao, Duy-Phuong, et al.
Pubblicazione: (2026)
FAR-Net: Multi-Stage Fusion Network with Enhanced Semantic Alignment and Adaptive Reconciliation for Composed Image Retrieval
di: Park, Jeong-Woo, et al.
Pubblicazione: (2025)
di: Park, Jeong-Woo, et al.
Pubblicazione: (2025)
Emotional Vietnamese Speech-Based Depression Diagnosis Using Dynamic Attention Mechanism
di: D., Quang-Anh N., et al.
Pubblicazione: (2024)
di: D., Quang-Anh N., et al.
Pubblicazione: (2024)
DEEPTalk: Dynamic Emotion Embedding for Probabilistic Speech-Driven 3D Face Animation
di: Kim, Jisoo, et al.
Pubblicazione: (2024)
di: Kim, Jisoo, et al.
Pubblicazione: (2024)
Tracking the Discriminative Axis: Dual Prototypes for Test-Time OOD Detection Under Covariate Shift
di: Lee, Wooseok, et al.
Pubblicazione: (2026)
di: Lee, Wooseok, et al.
Pubblicazione: (2026)
3D Prior is All You Need: Cross-Task Few-shot 2D Gaze Estimation
di: Cheng, Yihua, et al.
Pubblicazione: (2025)
di: Cheng, Yihua, et al.
Pubblicazione: (2025)
Leveraging Out-of-Distribution Unlabeled Images: Semi-Supervised Semantic Segmentation with an Open-Vocabulary Model
di: Shin, Wooseok, et al.
Pubblicazione: (2025)
di: Shin, Wooseok, et al.
Pubblicazione: (2025)
SenseShift6D: Multimodal RGB-D Benchmarking for Robust 6D Pose Estimation across Environment and Sensor Variations
di: Han, Yegyu, et al.
Pubblicazione: (2025)
di: Han, Yegyu, et al.
Pubblicazione: (2025)
Infinite-Homography as Robust Conditioning for Camera-Controlled Video Generation
di: Kim, Min-Jung, et al.
Pubblicazione: (2025)
di: Kim, Min-Jung, et al.
Pubblicazione: (2025)
Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling
di: Hyung, Junha, et al.
Pubblicazione: (2024)
di: Hyung, Junha, et al.
Pubblicazione: (2024)
State Estimation and Control of Dynamic Systems from High-Dimensional Image Data
di: Rasul, Ashik E, et al.
Pubblicazione: (2025)
di: Rasul, Ashik E, et al.
Pubblicazione: (2025)
Motion Cues from Image-based Point Tracking for LiDAR Scene Flow Estimation
di: Jang, Youngdong, et al.
Pubblicazione: (2026)
di: Jang, Youngdong, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Transformer with Leveraged Masked Autoencoder for video-based Pain Assessment
di: Nguyen, Minh-Duc, et al.
Pubblicazione: (2024) -
Anatomical Attention Alignment representation for Radiology Report Generation
di: Nguyen, Quang Vinh, et al.
Pubblicazione: (2025) -
ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise Diffusion
di: Vo, Hoang-Son, et al.
Pubblicazione: (2025) -
Latent Behavior Diffusion for Sequential Reaction Generation in Dyadic Setting
di: Nguyen, Minh-Duc, et al.
Pubblicazione: (2025) -
WaveNets: Wavelet Channel Attention Networks
di: Salman, Hadi, et al.
Pubblicazione: (2022)