InfinityHuman: Towards Long-Term Audio-Driven Human
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Xiaodi, Xie, Pan, Ren, Yi, Gan, Qijun, Zhang, Chen, Kong, Fangyuan, Yin, Xiang, Peng, Bingyue, Yuan, Zehuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
von: Han, Jian, et al.
Veröffentlicht: (2024)
von: Han, Jian, et al.
Veröffentlicht: (2024)
InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
von: Liu, Jinlai, et al.
Veröffentlicht: (2025)
von: Liu, Jinlai, et al.
Veröffentlicht: (2025)
ALIVE: Animate Your World with Lifelike Audio-Video Generation
von: Guo, Ying, et al.
Veröffentlicht: (2026)
von: Guo, Ying, et al.
Veröffentlicht: (2026)
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
von: Tian, Keyu, et al.
Veröffentlicht: (2024)
von: Tian, Keyu, et al.
Veröffentlicht: (2024)
Generative Refinement Networks for Visual Synthesis
von: Han, Jian, et al.
Veröffentlicht: (2026)
von: Han, Jian, et al.
Veröffentlicht: (2026)
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
von: Sun, Peize, et al.
Veröffentlicht: (2024)
von: Sun, Peize, et al.
Veröffentlicht: (2024)
UniTok: A Unified Tokenizer for Visual Generation and Understanding
von: Ma, Chuofan, et al.
Veröffentlicht: (2025)
von: Ma, Chuofan, et al.
Veröffentlicht: (2025)
MOSPA: Human Motion Generation Driven by Spatial Audio
von: Xu, Shuyang, et al.
Veröffentlicht: (2025)
von: Xu, Shuyang, et al.
Veröffentlicht: (2025)
VC-LLM: Automated Advertisement Video Creation from Raw Footage using Multi-modal LLMs
von: Qian, Dongjun, et al.
Veröffentlicht: (2025)
von: Qian, Dongjun, et al.
Veröffentlicht: (2025)
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
Waver: Wave Your Way to Lifelike Video Generation
von: Zhang, Yifu, et al.
Veröffentlicht: (2025)
von: Zhang, Yifu, et al.
Veröffentlicht: (2025)
FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
von: Zhang, Shilong, et al.
Veröffentlicht: (2025)
von: Zhang, Shilong, et al.
Veröffentlicht: (2025)
HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters
von: Chen, Yi, et al.
Veröffentlicht: (2025)
von: Chen, Yi, et al.
Veröffentlicht: (2025)
Unsupervised Discovery of Long-Term Spatiotemporal Periodic Workflows in Human Activities
von: Yang, Fan, et al.
Veröffentlicht: (2025)
von: Yang, Fan, et al.
Veröffentlicht: (2025)
Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency
von: Jiang, Jianwen, et al.
Veröffentlicht: (2024)
von: Jiang, Jianwen, et al.
Veröffentlicht: (2024)
From Parts to Whole: A Unified Reference Framework for Controllable Human Image Generation
von: Huang, Zehuan, et al.
Veröffentlicht: (2024)
von: Huang, Zehuan, et al.
Veröffentlicht: (2024)
TELA: Text to Layer-wise 3D Clothed Human Generation
von: Dong, Junting, et al.
Veröffentlicht: (2024)
von: Dong, Junting, et al.
Veröffentlicht: (2024)
Intention-Conditioned Long-Term Human Egocentric Action Forecasting
von: Mascaro, Esteve Valls, et al.
Veröffentlicht: (2022)
von: Mascaro, Esteve Valls, et al.
Veröffentlicht: (2022)
LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision
von: Li, Chunyu, et al.
Veröffentlicht: (2024)
von: Li, Chunyu, et al.
Veröffentlicht: (2024)
Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generation
von: Zhen, Dingcheng, et al.
Veröffentlicht: (2025)
von: Zhen, Dingcheng, et al.
Veröffentlicht: (2025)
AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation
von: Liang, Chao, et al.
Veröffentlicht: (2025)
von: Liang, Chao, et al.
Veröffentlicht: (2025)
Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length
von: Huang, Yubo, et al.
Veröffentlicht: (2025)
von: Huang, Yubo, et al.
Veröffentlicht: (2025)
PI-HMR: Towards Robust In-bed Temporal Human Shape Reconstruction with Contact Pressure Sensing
von: Wu, Ziyu, et al.
Veröffentlicht: (2025)
von: Wu, Ziyu, et al.
Veröffentlicht: (2025)
Fine-Grained Multi-View Hand Reconstruction Using Inverse Rendering
von: Gan, Qijun, et al.
Veröffentlicht: (2024)
von: Gan, Qijun, et al.
Veröffentlicht: (2024)
M3G: Multi-Granular Gesture Generator for Audio-Driven Full-Body Human Motion Synthesis
von: Yin, Zhizhuo, et al.
Veröffentlicht: (2025)
von: Yin, Zhizhuo, et al.
Veröffentlicht: (2025)
Video-Infinity: Distributed Long Video Generation
von: Tan, Zhenxiong, et al.
Veröffentlicht: (2024)
von: Tan, Zhenxiong, et al.
Veröffentlicht: (2024)
Lookahead Anchoring: Preserving Character Identity in Audio-Driven Human Animation
von: Seo, Junyoung, et al.
Veröffentlicht: (2025)
von: Seo, Junyoung, et al.
Veröffentlicht: (2025)
Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation
von: Li, Ruibin, et al.
Veröffentlicht: (2026)
von: Li, Ruibin, et al.
Veröffentlicht: (2026)
HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation
von: Huang, Ziyao, et al.
Veröffentlicht: (2025)
von: Huang, Ziyao, et al.
Veröffentlicht: (2025)
ITC-RWKV: Interactive Tissue-Cell Modeling with Recurrent Key-Value Aggregation for Histopathological Subtyping
von: Huang, Yating, et al.
Veröffentlicht: (2025)
von: Huang, Yating, et al.
Veröffentlicht: (2025)
Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked Autoregression
von: Meng, Zichong, et al.
Veröffentlicht: (2024)
von: Meng, Zichong, et al.
Veröffentlicht: (2024)
CamDirector: Towards Long-Term Coherent Video Trajectory Editing
von: Shi, Zhihao, et al.
Veröffentlicht: (2026)
von: Shi, Zhihao, et al.
Veröffentlicht: (2026)
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
von: Guan, Jiazhi, et al.
Veröffentlicht: (2025)
von: Guan, Jiazhi, et al.
Veröffentlicht: (2025)
LoSh: Long-Short Text Joint Prediction Network for Referring Video Object Segmentation
von: Yuan, Linfeng, et al.
Veröffentlicht: (2023)
von: Yuan, Linfeng, et al.
Veröffentlicht: (2023)
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
von: Wang, Junke, et al.
Veröffentlicht: (2024)
von: Wang, Junke, et al.
Veröffentlicht: (2024)
AudioScenic: Audio-Driven Video Scene Editing
von: Shen, Kaixin, et al.
Veröffentlicht: (2024)
von: Shen, Kaixin, et al.
Veröffentlicht: (2024)
T2LM: Long-Term 3D Human Motion Generation from Multiple Sentences
von: Lee, Taeryung, et al.
Veröffentlicht: (2024)
von: Lee, Taeryung, et al.
Veröffentlicht: (2024)
Generative Region-Language Pretraining for Open-Ended Object Detection
von: Lin, Chuang, et al.
Veröffentlicht: (2024)
von: Lin, Chuang, et al.
Veröffentlicht: (2024)
Stability-Driven Motion Generation for Object-Guided Human-Human Co-Manipulation
von: Xu, Jiahao, et al.
Veröffentlicht: (2026)
von: Xu, Jiahao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation
von: Gan, Qijun, et al.
Veröffentlicht: (2025) -
Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
von: Han, Jian, et al.
Veröffentlicht: (2024) -
InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
von: Liu, Jinlai, et al.
Veröffentlicht: (2025) -
ALIVE: Animate Your World with Lifelike Audio-Video Generation
von: Guo, Ying, et al.
Veröffentlicht: (2026) -
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
von: Tian, Keyu, et al.
Veröffentlicht: (2024)