INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Yongming, Zhang, Longhao, Rong, Zhengkun, Hu, Tianshu, Liang, Shuang, Ge, Zhipeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid Guidance
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
PersonaTalk: Bring Attention to Your Persona in Visual Dubbing
von: Zhang, Longhao, et al.
Veröffentlicht: (2024)
von: Zhang, Longhao, et al.
Veröffentlicht: (2024)
MobilePortrait: Real-Time One-Shot Neural Head Avatars on Mobile Devices
von: Jiang, Jianwen, et al.
Veröffentlicht: (2024)
von: Jiang, Jianwen, et al.
Veröffentlicht: (2024)
FlowAct-R1: Towards Interactive Humanoid Video Generation
von: Wang, Lizhen, et al.
Veröffentlicht: (2026)
von: Wang, Lizhen, et al.
Veröffentlicht: (2026)
DreamActor-M2: Universal Character Image Animation via Spatiotemporal In-Context Learning
von: Luo, Mingshuang, et al.
Veröffentlicht: (2026)
von: Luo, Mingshuang, et al.
Veröffentlicht: (2026)
OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions
von: Luo, Cheng, et al.
Veröffentlicht: (2025)
von: Luo, Cheng, et al.
Veröffentlicht: (2025)
Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2026)
Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language Models
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset
von: Agrawal, Vasu, et al.
Veröffentlicht: (2025)
von: Agrawal, Vasu, et al.
Veröffentlicht: (2025)
VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single Image
von: Xu, Sicheng, et al.
Veröffentlicht: (2025)
von: Xu, Sicheng, et al.
Veröffentlicht: (2025)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
von: Ling, Jun, et al.
Veröffentlicht: (2024)
von: Ling, Jun, et al.
Veröffentlicht: (2024)
Learning What to Trust: Bayesian Prior-Guided Optimization for Visual Generation
von: Liu, Ruiying, et al.
Veröffentlicht: (2025)
von: Liu, Ruiying, et al.
Veröffentlicht: (2025)
EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion
von: Wang, Haotian, et al.
Veröffentlicht: (2024)
von: Wang, Haotian, et al.
Veröffentlicht: (2024)
Category Query Learning for Human-Object Interaction Classification
von: Xie, Chi, et al.
Veröffentlicht: (2023)
von: Xie, Chi, et al.
Veröffentlicht: (2023)
PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation
von: Wang, Baiqin, et al.
Veröffentlicht: (2025)
von: Wang, Baiqin, et al.
Veröffentlicht: (2025)
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
von: Zhang, Mengchen, et al.
Veröffentlicht: (2025)
von: Zhang, Mengchen, et al.
Veröffentlicht: (2025)
Audio-Driven Talking Face Video Generation with Joint Uncertainty Learning
von: Xie, Yifan, et al.
Veröffentlicht: (2025)
von: Xie, Yifan, et al.
Veröffentlicht: (2025)
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent Space
von: Zhang, Zeren, et al.
Veröffentlicht: (2024)
von: Zhang, Zeren, et al.
Veröffentlicht: (2024)
Head Forcing: Long Autoregressive Video Generation via Head Heterogeneity
von: Tian, Jiahao, et al.
Veröffentlicht: (2026)
von: Tian, Jiahao, et al.
Veröffentlicht: (2026)
VidSketch: Hand-drawn Sketch-Driven Video Generation with Diffusion Control
von: Jiang, Lifan, et al.
Veröffentlicht: (2025)
von: Jiang, Lifan, et al.
Veröffentlicht: (2025)
Non-Markov Multi-Round Conversational Image Generation with History-Conditioned MLLMs
von: Zhang, Haochen, et al.
Veröffentlicht: (2026)
von: Zhang, Haochen, et al.
Veröffentlicht: (2026)
EmoGene: Audio-Driven Emotional 3D Talking-Head Generation
von: Wang, Wenqing, et al.
Veröffentlicht: (2024)
von: Wang, Wenqing, et al.
Veröffentlicht: (2024)
TempoSyncDiff: Distilled Temporally-Consistent Diffusion for Low-Latency Audio-Driven Talking Head Generation
von: Mazumdar, Soumya, et al.
Veröffentlicht: (2026)
von: Mazumdar, Soumya, et al.
Veröffentlicht: (2026)
UniHead: Unifying Multi-Perception for Detection Heads
von: Zhou, Hantao, et al.
Veröffentlicht: (2023)
von: Zhou, Hantao, et al.
Veröffentlicht: (2023)
FlowPalm: Optical Flow Driven Non-Rigid Deformation for Geometrically Diverse Palmprint Generation
von: Zou, Yuchen, et al.
Veröffentlicht: (2026)
von: Zou, Yuchen, et al.
Veröffentlicht: (2026)
BREEN: Bridge Data-Efficient Encoder-Free Multimodal Learning with Learnable Queries
von: Li, Tianle, et al.
Veröffentlicht: (2025)
von: Li, Tianle, et al.
Veröffentlicht: (2025)
Superior and Pragmatic Talking Face Generation with Teacher-Student Framework
von: Liang, Chao, et al.
Veröffentlicht: (2024)
von: Liang, Chao, et al.
Veröffentlicht: (2024)
DeepFake Detection in Dyadic Video Calls using Point of Gaze Tracking
von: Kohler, Odin, et al.
Veröffentlicht: (2025)
von: Kohler, Odin, et al.
Veröffentlicht: (2025)
Empathetic Response in Audio-Visual Conversations Using Emotion Preference Optimization and MambaCompressor
von: Kim, Yeonju, et al.
Veröffentlicht: (2024)
von: Kim, Yeonju, et al.
Veröffentlicht: (2024)
Dynamic Multi-Target Fusion for Efficient Audio-Visual Navigation
von: Yu, Yinfeng, et al.
Veröffentlicht: (2025)
von: Yu, Yinfeng, et al.
Veröffentlicht: (2025)
Interact-Custom: Customized Human Object Interaction Image Generation
von: Xu, Zhu, et al.
Veröffentlicht: (2025)
von: Xu, Zhu, et al.
Veröffentlicht: (2025)
DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers
von: Wang, Lizhen, et al.
Veröffentlicht: (2025)
von: Wang, Lizhen, et al.
Veröffentlicht: (2025)
Consistent Video Editing as Flow-Driven Image-to-Video Generation
von: Wang, Ge, et al.
Veröffentlicht: (2025)
von: Wang, Ge, et al.
Veröffentlicht: (2025)
Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation
von: Ki, Taekyung, et al.
Veröffentlicht: (2026)
von: Ki, Taekyung, et al.
Veröffentlicht: (2026)
ExGes: Expressive Human Motion Retrieval and Modulation for Audio-Driven Gesture Synthesis
von: Zhou, Xukun, et al.
Veröffentlicht: (2025)
von: Zhou, Xukun, et al.
Veröffentlicht: (2025)
Observation-Aligned Mask Priors for Learning Physical Dynamics from Authentic Occlusions
von: Ma, Chiyuan, et al.
Veröffentlicht: (2026)
von: Ma, Chiyuan, et al.
Veröffentlicht: (2026)
Exploring Audio Hallucination in Egocentric Video Understanding
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2025)
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2025)
EAvatar: Expression-Aware Head Avatar Reconstruction with Generative Geometry Priors
von: Zhang, Shikun, et al.
Veröffentlicht: (2025)
von: Zhang, Shikun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid Guidance
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025) -
PersonaTalk: Bring Attention to Your Persona in Visual Dubbing
von: Zhang, Longhao, et al.
Veröffentlicht: (2024) -
MobilePortrait: Real-Time One-Shot Neural Head Avatars on Mobile Devices
von: Jiang, Jianwen, et al.
Veröffentlicht: (2024) -
FlowAct-R1: Towards Interactive Humanoid Video Generation
von: Wang, Lizhen, et al.
Veröffentlicht: (2026) -
DreamActor-M2: Universal Character Image Animation via Spatiotemporal In-Context Learning
von: Luo, Mingshuang, et al.
Veröffentlicht: (2026)