KAN-Based Fusion of Dual-Domain for Audio-Driven Facial Landmarks Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Vo-Thanh, Hoang-Son, Nguyen, Quang-Vinh, Kim, Soo-Hyung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise Diffusion
di: Vo, Hoang-Son, et al.
Pubblicazione: (2025)
di: Vo, Hoang-Son, et al.
Pubblicazione: (2025)
Talking Head Generation Driven by Speech-Related Facial Action Units and Audio- Based on Multimodal Representation Fusion
di: Chen, Sen, et al.
Pubblicazione: (2022)
di: Chen, Sen, et al.
Pubblicazione: (2022)
Anatomical Attention Alignment representation for Radiology Report Generation
di: Nguyen, Quang Vinh, et al.
Pubblicazione: (2025)
di: Nguyen, Quang Vinh, et al.
Pubblicazione: (2025)
Instruction-Driven 3D Facial Expression Generation and Transition
di: Vo, Anh H., et al.
Pubblicazione: (2026)
di: Vo, Anh H., et al.
Pubblicazione: (2026)
MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning
di: Truong, Quang-Trung, et al.
Pubblicazione: (2025)
di: Truong, Quang-Trung, et al.
Pubblicazione: (2025)
Rethinking Top Probability from Multi-view for Distracted Driver Behaviour Localization
di: Nguyen, Quang Vinh, et al.
Pubblicazione: (2024)
di: Nguyen, Quang Vinh, et al.
Pubblicazione: (2024)
Neural B-Frame Coding: Tackling Domain Shift Issues with Lightweight Online Motion Resolution Adaptation
di: NguyenQuang, Sang, et al.
Pubblicazione: (2025)
di: NguyenQuang, Sang, et al.
Pubblicazione: (2025)
Boosting Facial Action Unit Detection Through Jointly Learning Facial Landmark Detection and Domain Separation and Reconstruction
di: Shang, Ziqiao, et al.
Pubblicazione: (2023)
di: Shang, Ziqiao, et al.
Pubblicazione: (2023)
Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models
di: Jin, Hyundong, et al.
Pubblicazione: (2025)
di: Jin, Hyundong, et al.
Pubblicazione: (2025)
SFFNet: Synergistic Feature Fusion Network With Dual-Domain Edge Enhancement for UAV Image Object Detection
di: Zhang, Wenfeng, et al.
Pubblicazione: (2026)
di: Zhang, Wenfeng, et al.
Pubblicazione: (2026)
ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer
di: Guan, Jiazhi, et al.
Pubblicazione: (2024)
di: Guan, Jiazhi, et al.
Pubblicazione: (2024)
AdaMesh: Personalized Facial Expressions and Head Poses for Adaptive Speech-Driven 3D Facial Animation
di: Chen, Liyang, et al.
Pubblicazione: (2023)
di: Chen, Liyang, et al.
Pubblicazione: (2023)
A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation
di: Zhou, S. Z., et al.
Pubblicazione: (2025)
di: Zhou, S. Z., et al.
Pubblicazione: (2025)
CONSTANT: Towards High-Quality One-Shot Handwriting Generation with Patch Contrastive Enhancement and Style-Aware Quantization
di: Le, Anh-Duy, et al.
Pubblicazione: (2026)
di: Le, Anh-Duy, et al.
Pubblicazione: (2026)
A Dual-Module Denoising Approach with Curriculum Learning for Enhancing Multimodal Aspect-Based Sentiment Analysis
di: Van Doan, Nguyen, et al.
Pubblicazione: (2024)
di: Van Doan, Nguyen, et al.
Pubblicazione: (2024)
ReactDiff: Latent Diffusion for Facial Reaction Generation
di: Li, Jiaming, et al.
Pubblicazione: (2025)
di: Li, Jiaming, et al.
Pubblicazione: (2025)
Multi-source Multimodal Progressive Domain Adaption for Audio-Visual Deception Detection
di: Lin, Ronghao, et al.
Pubblicazione: (2025)
di: Lin, Ronghao, et al.
Pubblicazione: (2025)
KeyNode-Driven Geometry Coding for Real-World Scanned Human Dynamic Mesh Compression
di: Hoang, Huong, et al.
Pubblicazione: (2025)
di: Hoang, Huong, et al.
Pubblicazione: (2025)
GaussianTalker: Real-Time High-Fidelity Talking Head Synthesis with Audio-Driven 3D Gaussian Splatting
di: Cho, Kyusun, et al.
Pubblicazione: (2024)
di: Cho, Kyusun, et al.
Pubblicazione: (2024)
Polyp-SES: Automatic Polyp Segmentation with Self-Enriched Semantic Model
di: Nguyen, Quang Vinh, et al.
Pubblicazione: (2024)
di: Nguyen, Quang Vinh, et al.
Pubblicazione: (2024)
Towards Signboard-Oriented Visual Question Answering: ViSignVQA Dataset, Method and Benchmark
di: Nguyen, Hieu Minh, et al.
Pubblicazione: (2025)
di: Nguyen, Hieu Minh, et al.
Pubblicazione: (2025)
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
di: Guan, Jiazhi, et al.
Pubblicazione: (2025)
di: Guan, Jiazhi, et al.
Pubblicazione: (2025)
EditEmoTalk: Controllable Speech-Driven 3D Facial Animation with Continuous Expression Editing
di: Jiang, Diqiong, et al.
Pubblicazione: (2026)
di: Jiang, Diqiong, et al.
Pubblicazione: (2026)
FD2Talk: Towards Generalized Talking Head Generation with Facial Decoupled Diffusion Model
di: Yao, Ziyu, et al.
Pubblicazione: (2024)
di: Yao, Ziyu, et al.
Pubblicazione: (2024)
TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation
di: Liu, Xiangyu, et al.
Pubblicazione: (2026)
di: Liu, Xiangyu, et al.
Pubblicazione: (2026)
KAN-SAM: Kolmogorov-Arnold Network Guided Segment Anything Model for RGB-T Salient Object Detection
di: Li, Xingyuan, et al.
Pubblicazione: (2025)
di: Li, Xingyuan, et al.
Pubblicazione: (2025)
DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization
di: Nguyen, Ngoc-Son, et al.
Pubblicazione: (2026)
di: Nguyen, Ngoc-Son, et al.
Pubblicazione: (2026)
MoLT: Mixture of Layer-Wise Tokens for Efficient Audio-Visual Learning
di: Rho, Kyeongha, et al.
Pubblicazione: (2025)
di: Rho, Kyeongha, et al.
Pubblicazione: (2025)
WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object Detection
di: Zhu, Haodong, et al.
Pubblicazione: (2025)
di: Zhu, Haodong, et al.
Pubblicazione: (2025)
TA-V2A: Textually Assisted Video-to-Audio Generation
di: You, Yuhuan, et al.
Pubblicazione: (2025)
di: You, Yuhuan, et al.
Pubblicazione: (2025)
EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion Transformers
di: Flynn, John, et al.
Pubblicazione: (2026)
di: Flynn, John, et al.
Pubblicazione: (2026)
A Sleep Monitoring System Based on Audio, Video and Depth Information
di: Chen, Lyn Chao-ling, et al.
Pubblicazione: (2025)
di: Chen, Lyn Chao-ling, et al.
Pubblicazione: (2025)
Dynamic Resolution Guidance for Facial Expression Recognition
di: Wang, Songpan, et al.
Pubblicazione: (2024)
di: Wang, Songpan, et al.
Pubblicazione: (2024)
AudioStory: Generating Long-Form Narrative Audio with Large Language Models
di: Guo, Yuxin, et al.
Pubblicazione: (2025)
di: Guo, Yuxin, et al.
Pubblicazione: (2025)
Bridging Your Imagination with Audio-Video Generation via a Unified Director
di: Zhang, Jiaxu, et al.
Pubblicazione: (2025)
di: Zhang, Jiaxu, et al.
Pubblicazione: (2025)
MLLM-VADStory: Domain Knowledge-Driven Multimodal LLMs for Video Ad Storyline Insights
di: Yang, Jasmine, et al.
Pubblicazione: (2026)
di: Yang, Jasmine, et al.
Pubblicazione: (2026)
Optimized Learned Image Compression for Facial Expression Recognition
di: Li, Xiumei, et al.
Pubblicazione: (2025)
di: Li, Xiumei, et al.
Pubblicazione: (2025)
Multimodal Engagement Analysis from Facial Videos in the Classroom
di: Sümer, Ömer, et al.
Pubblicazione: (2021)
di: Sümer, Ömer, et al.
Pubblicazione: (2021)
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
di: Pu, Junfu, et al.
Pubblicazione: (2026)
di: Pu, Junfu, et al.
Pubblicazione: (2026)
Federated Prompt-Tuning with Heterogeneous and Incomplete Multimodal Client Data
di: Phung, Thu Hang, et al.
Pubblicazione: (2026)
di: Phung, Thu Hang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise Diffusion
di: Vo, Hoang-Son, et al.
Pubblicazione: (2025) -
Talking Head Generation Driven by Speech-Related Facial Action Units and Audio- Based on Multimodal Representation Fusion
di: Chen, Sen, et al.
Pubblicazione: (2022) -
Anatomical Attention Alignment representation for Radiology Report Generation
di: Nguyen, Quang Vinh, et al.
Pubblicazione: (2025) -
Instruction-Driven 3D Facial Expression Generation and Transition
di: Vo, Anh H., et al.
Pubblicazione: (2026) -
MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning
di: Truong, Quang-Trung, et al.
Pubblicazione: (2025)