From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations
Fuente:
arXiv
Saved in:
| Main Authors: | Ng, Evonne, Romero, Javier, Bagautdinov, Timur, Bai, Shaojie, Darrell, Trevor, Kanazawa, Angjoo, Richard, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diffusion Forcing for Multi-Agent Interaction Sequence Modeling
by: Maluleke, Vongani H., et al.
Published: (2025)
by: Maluleke, Vongani H., et al.
Published: (2025)
Human-level 3D shape perception emerges from multi-view learning
by: Bonnen, Tyler, et al.
Published: (2026)
by: Bonnen, Tyler, et al.
Published: (2026)
Pose Priors from Language Models
by: Subramanian, Sanjay, et al.
Published: (2024)
by: Subramanian, Sanjay, et al.
Published: (2024)
SARAH: Spatially Aware Real-time Agentic Humans
by: Ng, Evonne, et al.
Published: (2026)
by: Ng, Evonne, et al.
Published: (2026)
Drivable 3D Gaussian Avatars
by: Zielonka, Wojciech, et al.
Published: (2023)
by: Zielonka, Wojciech, et al.
Published: (2023)
St4RTrack: Simultaneous 4D Reconstruction and Tracking in the World
by: Feng, Haiwen, et al.
Published: (2025)
by: Feng, Haiwen, et al.
Published: (2025)
Generating Continual Human Motion in Diverse 3D Scenes
by: Mir, Aymen, et al.
Published: (2023)
by: Mir, Aymen, et al.
Published: (2023)
Splatfacto-W: A Nerfstudio Implementation of Gaussian Splatting for Unconstrained Photo Collections
by: Xu, Congrong, et al.
Published: (2024)
by: Xu, Congrong, et al.
Published: (2024)
TurboPortrait3D: Single-step diffusion-based fast portrait novel-view synthesis
by: Kim, Emily, et al.
Published: (2025)
by: Kim, Emily, et al.
Published: (2025)
SOAR: Self-Occluded Avatar Recovery from a Single Video In the Wild
by: Pan, Zhuoyang, et al.
Published: (2024)
by: Pan, Zhuoyang, et al.
Published: (2024)
ASH: Animatable Gaussian Splats for Efficient and Photoreal Human Rendering
by: Pang, Haokai, et al.
Published: (2023)
by: Pang, Haokai, et al.
Published: (2023)
Sapiens: Foundation for Human Vision Models
by: Khirodkar, Rawal, et al.
Published: (2024)
by: Khirodkar, Rawal, et al.
Published: (2024)
Visually Prompted Benchmarks Are Surprisingly Fragile
by: Feng, Haiwen, et al.
Published: (2025)
by: Feng, Haiwen, et al.
Published: (2025)
Talking Together: Synthesizing Co-Located 3D Conversations from Audio
by: Shan, Mengyi, et al.
Published: (2026)
by: Shan, Mengyi, et al.
Published: (2026)
Visual Imitation Enables Contextual Humanoid Control
by: Allshire, Arthur, et al.
Published: (2025)
by: Allshire, Arthur, et al.
Published: (2025)
EgoRelight: Egocentric Human Capture and Illumination Recovery for Relightable and Photoreal Avatar Rendering
by: Chen, Jianchun, et al.
Published: (2026)
by: Chen, Jianchun, et al.
Published: (2026)
RePerformer: Immersive Human-centric Volumetric Videos from Playback to Photoreal Reperformance
by: Jiang, Yuheng, et al.
Published: (2025)
by: Jiang, Yuheng, et al.
Published: (2025)
The More You See in 2D, the More You Perceive in 3D
by: Han, Xinyang, et al.
Published: (2024)
by: Han, Xinyang, et al.
Published: (2024)
Agent-to-Sim: Learning Interactive Behavior Models from Casual Longitudinal Videos
by: Yang, Gengshan, et al.
Published: (2024)
by: Yang, Gengshan, et al.
Published: (2024)
GeoRelight: Learning Joint Geometrical Relighting and Reconstruction with Flexible Multi-Modal Diffusion Transformers
by: Xue, Yuxuan, et al.
Published: (2026)
by: Xue, Yuxuan, et al.
Published: (2026)
TEDRA: Text-based Editing of Dynamic and Photoreal Actors
by: Sunagad, Basavaraj, et al.
Published: (2024)
by: Sunagad, Basavaraj, et al.
Published: (2024)
Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
by: Yin, Shaofeng, et al.
Published: (2026)
by: Yin, Shaofeng, et al.
Published: (2026)
Continuous 3D Perception Model with Persistent State
by: Wang, Qianqian, et al.
Published: (2025)
by: Wang, Qianqian, et al.
Published: (2025)
NeRF-XL: Scaling NeRFs with Multiple GPUs
by: Li, Ruilong, et al.
Published: (2024)
by: Li, Ruilong, et al.
Published: (2024)
Reconstructing People, Places, and Cameras
by: Müller, Lea, et al.
Published: (2024)
by: Müller, Lea, et al.
Published: (2024)
Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind
by: Plizzari, Chiara, et al.
Published: (2024)
by: Plizzari, Chiara, et al.
Published: (2024)
Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving
by: Wang, Jiahao, et al.
Published: (2026)
by: Wang, Jiahao, et al.
Published: (2026)
DuoMo: Dual Motion Diffusion for World-Space Human Reconstruction
by: Wang, Yufu, et al.
Published: (2026)
by: Wang, Yufu, et al.
Published: (2026)
Pippo: High-Resolution Multi-View Humans from a Single Image
by: Kant, Yash, et al.
Published: (2025)
by: Kant, Yash, et al.
Published: (2025)
MAtCha Gaussians: Atlas of Charts for High-Quality Geometry and Photorealism From Sparse Views
by: Guédon, Antoine, et al.
Published: (2024)
by: Guédon, Antoine, et al.
Published: (2024)
Vector Quantized Feature Fields for Fast 3D Semantic Lifting
by: Tang, George, et al.
Published: (2025)
by: Tang, George, et al.
Published: (2025)
Finding Visual Task Vectors
by: Hojel, Alberto, et al.
Published: (2024)
by: Hojel, Alberto, et al.
Published: (2024)
Toon3D: Seeing Cartoons from New Perspectives
by: Weber, Ethan, et al.
Published: (2024)
by: Weber, Ethan, et al.
Published: (2024)
Photoreal Scene Reconstruction from an Egocentric Device
by: Lv, Zhaoyang, et al.
Published: (2025)
by: Lv, Zhaoyang, et al.
Published: (2025)
Segment Any Motion in Videos
by: Huang, Nan, et al.
Published: (2025)
by: Huang, Nan, et al.
Published: (2025)
Shape of Motion: 4D Reconstruction from a Single Video
by: Wang, Qianqian, et al.
Published: (2024)
by: Wang, Qianqian, et al.
Published: (2024)
Synergy and Synchrony in Couple Dances
by: Maluleke, Vongani, et al.
Published: (2024)
by: Maluleke, Vongani, et al.
Published: (2024)
MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos
by: Li, Zhengqi, et al.
Published: (2024)
by: Li, Zhengqi, et al.
Published: (2024)
From Generated Human Videos to Physically Plausible Robot Trajectories
by: Ni, James, et al.
Published: (2025)
by: Ni, James, et al.
Published: (2025)
Decentralized Diffusion Models
by: McAllister, David, et al.
Published: (2025)
by: McAllister, David, et al.
Published: (2025)
Similar Items
-
Diffusion Forcing for Multi-Agent Interaction Sequence Modeling
by: Maluleke, Vongani H., et al.
Published: (2025) -
Human-level 3D shape perception emerges from multi-view learning
by: Bonnen, Tyler, et al.
Published: (2026) -
Pose Priors from Language Models
by: Subramanian, Sanjay, et al.
Published: (2024) -
SARAH: Spatially Aware Real-time Agentic Humans
by: Ng, Evonne, et al.
Published: (2026) -
Drivable 3D Gaussian Avatars
by: Zielonka, Wojciech, et al.
Published: (2023)