PersonaTalk: Bring Attention to Your Persona in Visual Dubbing
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Longhao, Liang, Shuang, Ge, Zhipeng, Hu, Tianshu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations
by: Zhu, Yongming, et al.
Published: (2024)
by: Zhu, Yongming, et al.
Published: (2024)
Dubbing for Everyone: Data-Efficient Visual Dubbing using Neural Rendering Priors
by: Saunders, Jack, et al.
Published: (2024)
by: Saunders, Jack, et al.
Published: (2024)
Learning Disentangled Speech- and Expression-Driven Blendshapes for 3D Talking Face Animation
by: Mao, Yuxiang, et al.
Published: (2025)
by: Mao, Yuxiang, et al.
Published: (2025)
Lightning Fast Caching-based Parallel Denoising Prediction for Accelerating Talking Head Generation
by: Long, Jianzhi, et al.
Published: (2025)
by: Long, Jianzhi, et al.
Published: (2025)
Intentional Gesture: Deliver Your Intentions with Gestures for Speech
by: Liu, Pinxin, et al.
Published: (2025)
by: Liu, Pinxin, et al.
Published: (2025)
JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion
by: Chen, Anthony, et al.
Published: (2026)
by: Chen, Anthony, et al.
Published: (2026)
VideoPanda: Video Panoramic Diffusion with Multi-view Attention
by: Xie, Kevin, et al.
Published: (2025)
by: Xie, Kevin, et al.
Published: (2025)
MoA: Mixture-of-Attention for Subject-Context Disentanglement in Personalized Image Generation
by: Wang, Kuan-Chieh, et al.
Published: (2024)
by: Wang, Kuan-Chieh, et al.
Published: (2024)
Self-Attention Based Multi-Scale Graph Auto-Encoder Network of 3D Meshes
by: Nazir, Saqib, et al.
Published: (2025)
by: Nazir, Saqib, et al.
Published: (2025)
Object-level Visual Prompts for Compositional Image Generation
by: Parmar, Gaurav, et al.
Published: (2025)
by: Parmar, Gaurav, et al.
Published: (2025)
Generating by Understanding: Neural Visual Generation with Logical Symbol Groundings
by: Peng, Yifei, et al.
Published: (2023)
by: Peng, Yifei, et al.
Published: (2023)
PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion Representation
by: Zhao, Junchuan, et al.
Published: (2026)
by: Zhao, Junchuan, et al.
Published: (2026)
StyleRF-VolVis: Style Transfer of Neural Radiance Fields for Expressive Volume Visualization
by: Tang, Kaiyuan, et al.
Published: (2024)
by: Tang, Kaiyuan, et al.
Published: (2024)
Progressive Rendering Distillation: Adapting Stable Diffusion for Instant Text-to-Mesh Generation without 3D Data
by: Ma, Zhiyuan, et al.
Published: (2025)
by: Ma, Zhiyuan, et al.
Published: (2025)
Bringing Attention to CAD: Boundary Representation Learning via Transformer
by: Zou, Qiang, et al.
Published: (2025)
by: Zou, Qiang, et al.
Published: (2025)
Ancestral Mamba: Enhancing Selective Discriminant Space Model with Online Visual Prototype Learning for Efficient and Robust Discriminant Approach
by: Qin, Jiahao, et al.
Published: (2025)
by: Qin, Jiahao, et al.
Published: (2025)
A 3D Generation Framework from Cross Modality to Parameterized Primitive
by: Liang, Yiming, et al.
Published: (2025)
by: Liang, Yiming, et al.
Published: (2025)
Synthesizing Physically Plausible Human Motions in 3D Scenes
by: Pan, Liang, et al.
Published: (2023)
by: Pan, Liang, et al.
Published: (2023)
GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors
by: Xu, Tian-Xing, et al.
Published: (2025)
by: Xu, Tian-Xing, et al.
Published: (2025)
DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
by: Hu, Wenbo, et al.
Published: (2024)
by: Hu, Wenbo, et al.
Published: (2024)
Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
by: Yin, Shaofeng, et al.
Published: (2026)
by: Yin, Shaofeng, et al.
Published: (2026)
TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models
by: YU, Mark, et al.
Published: (2025)
by: YU, Mark, et al.
Published: (2025)
HY-Motion 1.0: Scaling Flow Matching Models for Text-To-Motion Generation
by: Wen, Yuxin, et al.
Published: (2025)
by: Wen, Yuxin, et al.
Published: (2025)
LuxDiT: Lighting Estimation with Video Diffusion Transformer
by: Liang, Ruofan, et al.
Published: (2025)
by: Liang, Ruofan, et al.
Published: (2025)
Photorealistic Object Insertion with Diffusion-Guided Inverse Rendering
by: Liang, Ruofan, et al.
Published: (2024)
by: Liang, Ruofan, et al.
Published: (2024)
Lodge++: High-quality and Long Dance Generation with Vivid Choreography Patterns
by: Li, Ronghui, et al.
Published: (2024)
by: Li, Ronghui, et al.
Published: (2024)
SeqTex: Generate Mesh Textures in Video Sequence
by: Yuan, Ze, et al.
Published: (2025)
by: Yuan, Ze, et al.
Published: (2025)
SCube: Instant Large-Scale Scene Reconstruction using VoxSplats
by: Ren, Xuanchi, et al.
Published: (2024)
by: Ren, Xuanchi, et al.
Published: (2024)
One Shot, One Talk: Whole-body Talking Avatar from a Single Image
by: Xiang, Jun, et al.
Published: (2024)
by: Xiang, Jun, et al.
Published: (2024)
Edify 3D: Scalable High-Quality 3D Asset Generation
by: NVIDIA, et al.
Published: (2024)
by: NVIDIA, et al.
Published: (2024)
Feed-Forward Bullet-Time Reconstruction of Dynamic Scenes from Monocular Videos
by: Liang, Hanxue, et al.
Published: (2024)
by: Liang, Hanxue, et al.
Published: (2024)
TEXGen: a Generative Diffusion Model for Mesh Textures
by: Yu, Xin, et al.
Published: (2024)
by: Yu, Xin, et al.
Published: (2024)
Learning to Synthesize Graphics Programs for Geometric Artworks
by: Bing, Qi, et al.
Published: (2024)
by: Bing, Qi, et al.
Published: (2024)
OT-Talk: Animating 3D Talking Head with Optimal Transportation
by: Wang, Xinmu, et al.
Published: (2025)
by: Wang, Xinmu, et al.
Published: (2025)
OmniHands: Towards Robust 4D Hand Mesh Recovery via A Versatile Transformer
by: Lin, Dixuan, et al.
Published: (2024)
by: Lin, Dixuan, et al.
Published: (2024)
Mojito: LLM-Aided Motion Instructor with Jitter-Reduced Inertial Tokens
by: Shan, Ziwei, et al.
Published: (2025)
by: Shan, Ziwei, et al.
Published: (2025)
CoARF: Controllable 3D Artistic Style Transfer for Radiance Fields
by: Zhang, Deheng, et al.
Published: (2024)
by: Zhang, Deheng, et al.
Published: (2024)
VOODOO XP: Expressive One-Shot Head Reenactment for VR Telepresence
by: Tran, Phong, et al.
Published: (2024)
by: Tran, Phong, et al.
Published: (2024)
Learning to Edit Visual Programs with Self-Supervision
by: Jones, R. Kenny, et al.
Published: (2024)
by: Jones, R. Kenny, et al.
Published: (2024)
Light-SQ: Structure-aware Shape Abstraction with Superquadrics for Generated Meshes
by: Wang, Yuhan, et al.
Published: (2025)
by: Wang, Yuhan, et al.
Published: (2025)
Similar Items
-
INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations
by: Zhu, Yongming, et al.
Published: (2024) -
Dubbing for Everyone: Data-Efficient Visual Dubbing using Neural Rendering Priors
by: Saunders, Jack, et al.
Published: (2024) -
Learning Disentangled Speech- and Expression-Driven Blendshapes for 3D Talking Face Animation
by: Mao, Yuxiang, et al.
Published: (2025) -
Lightning Fast Caching-based Parallel Denoising Prediction for Accelerating Talking Head Generation
by: Long, Jianzhi, et al.
Published: (2025) -
Intentional Gesture: Deliver Your Intentions with Gestures for Speech
by: Liu, Pinxin, et al.
Published: (2025)