InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zhenzhi, Yang, Jiaqi, Jiang, Jianwen, Liang, Chao, Lin, Gaojie, Zheng, Zerong, Yang, Ceyuan, Zhang, Yuan, Gao, Mingyuan, Lin, Dahua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models
by: Lin, Gaojie, et al.
Published: (2025)
by: Lin, Gaojie, et al.
Published: (2025)
AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation
by: Liang, Chao, et al.
Published: (2025)
by: Liang, Chao, et al.
Published: (2025)
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
by: Wang, Zhenzhi, et al.
Published: (2025)
by: Wang, Zhenzhi, et al.
Published: (2025)
FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation
by: Zhong, Tianyun, et al.
Published: (2024)
by: Zhong, Tianyun, et al.
Published: (2024)
Semantics-Aware Human Motion Generation from Audio Instructions
by: Wang, Zi-An, et al.
Published: (2025)
by: Wang, Zi-An, et al.
Published: (2025)
BATON: Aligning Text-to-Audio Model with Human Preference Feedback
by: Liao, Huan, et al.
Published: (2024)
by: Liao, Huan, et al.
Published: (2024)
Aligning Audio Captions with Human Preferences
by: Hegde, Kartik, et al.
Published: (2025)
by: Hegde, Kartik, et al.
Published: (2025)
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models
by: Yang, Jialiang, et al.
Published: (2026)
by: Yang, Jialiang, et al.
Published: (2026)
UniAudio 2.0: A Unified Audio Language Model with Text-Aligned Factorized Audio Tokenization
by: Yang, Dongchao, et al.
Published: (2026)
by: Yang, Dongchao, et al.
Published: (2026)
OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
by: Jiang, Jianwen, et al.
Published: (2025)
by: Jiang, Jianwen, et al.
Published: (2025)
AlignCap: Aligning Speech Emotion Captioning to Human Preferences
by: Liang, Ziqi, et al.
Published: (2024)
by: Liang, Ziqi, et al.
Published: (2024)
Semantic Proximity Alignment: Towards Human Perception-consistent Audio Tagging by Aligning with Label Text Description
by: Liu, Wuyang, et al.
Published: (2023)
by: Liu, Wuyang, et al.
Published: (2023)
InterControl: Zero-shot Human Interaction Generation by Controlling Every Joint
by: Wang, Zhenzhi, et al.
Published: (2023)
by: Wang, Zhenzhi, et al.
Published: (2023)
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
by: Liu, Qianhui, et al.
Published: (2024)
by: Liu, Qianhui, et al.
Published: (2024)
STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence
by: Liu, Zihan, et al.
Published: (2025)
by: Liu, Zihan, et al.
Published: (2025)
ERIS: Evolutionary Real-world Interference Scheme for Jailbreaking Audio Large Models
by: Zhang, Yibo, et al.
Published: (2025)
by: Zhang, Yibo, et al.
Published: (2025)
DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners
by: Luo, Xiaoxue, et al.
Published: (2025)
by: Luo, Xiaoxue, et al.
Published: (2025)
Towards Automatic Evaluation and High-Quality Pseudo-Parallel Dataset Construction for Audio Editing: A Human-in-the-Loop Method
by: Jia, Yuhang, et al.
Published: (2025)
by: Jia, Yuhang, et al.
Published: (2025)
Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency
by: Jiang, Jianwen, et al.
Published: (2024)
by: Jiang, Jianwen, et al.
Published: (2024)
CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention
by: Lin, Gaojie, et al.
Published: (2024)
by: Lin, Gaojie, et al.
Published: (2024)
FoleySpace: Vision-Aligned Binaural Spatial Audio Generation
by: Zhao, Lei, et al.
Published: (2025)
by: Zhao, Lei, et al.
Published: (2025)
ChronosAudio: A Comprehensive Long-Audio Benchmark for Evaluating Audio-Large Language Models
by: Luo, Kaiwen, et al.
Published: (2026)
by: Luo, Kaiwen, et al.
Published: (2026)
Sonic: Shifting Focus to Global Audio Perception in Portrait Animation
by: Ji, Xiaozhong, et al.
Published: (2024)
by: Ji, Xiaozhong, et al.
Published: (2024)
SpeechAlign: Aligning Speech Generation to Human Preferences
by: Zhang, Dong, et al.
Published: (2024)
by: Zhang, Dong, et al.
Published: (2024)
Multi-identity Human Image Animation with Structural Video Diffusion
by: Wang, Zhenzhi, et al.
Published: (2025)
by: Wang, Zhenzhi, et al.
Published: (2025)
Content and Style Aware Audio-Driven Facial Animation
by: Liu, Qingju, et al.
Published: (2024)
by: Liu, Qingju, et al.
Published: (2024)
Correlation of Fréchet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant
by: Tailleur, Modan, et al.
Published: (2024)
by: Tailleur, Modan, et al.
Published: (2024)
APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track
by: Miao, Deshui, et al.
Published: (2026)
by: Miao, Deshui, et al.
Published: (2026)
Seconds-Aligned PCA-DAC Latent Diffusion for Symbolic-to-Audio Drum Rendering
by: Soiledis, Konstantinos, et al.
Published: (2026)
by: Soiledis, Konstantinos, et al.
Published: (2026)
DualMark: Identifying Model and Training Data Origins in Generated Audio
by: Yang, Xuefeng, et al.
Published: (2025)
by: Yang, Xuefeng, et al.
Published: (2025)
GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human
by: Lin, Yihang, et al.
Published: (2026)
by: Lin, Yihang, et al.
Published: (2026)
Continual Audio Deepfake Detection via Universal Adversarial Perturbation
by: Li, Wangjie, et al.
Published: (2025)
by: Li, Wangjie, et al.
Published: (2025)
Aligning Text-to-Music Evaluation with Human Preferences
by: Huang, Yichen, et al.
Published: (2025)
by: Huang, Yichen, et al.
Published: (2025)
Speech-DRAME: A Framework for Human-Aligned Benchmarks in Speech Role-Play
by: Shi, Jiatong, et al.
Published: (2025)
by: Shi, Jiatong, et al.
Published: (2025)
RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer
by: Du, Fangyu, et al.
Published: (2025)
by: Du, Fangyu, et al.
Published: (2025)
AR&D: A Framework for Retrieving and Describing Concepts for Interpreting AudioLLMs
by: Chowdhury, Townim Faisal, et al.
Published: (2026)
by: Chowdhury, Townim Faisal, et al.
Published: (2026)
Contrastive Learning With Audio Discrimination For Customizable Keyword Spotting In Continuous Speech
by: Xi, Yu, et al.
Published: (2024)
by: Xi, Yu, et al.
Published: (2024)
MoEScore: Mixture-of-Experts-Based Text-Audio Relevance Score Prediction for Text-to-Audio System Evaluation
by: Sun, Bochao, et al.
Published: (2026)
by: Sun, Bochao, et al.
Published: (2026)
ByteComposer: a Human-like Melody Composition Method based on Language Model Agent
by: Liang, Xia, et al.
Published: (2024)
by: Liang, Xia, et al.
Published: (2024)
Audio Mamba: Pretrained Audio State Space Model For Audio Tagging
by: Lin, Jiaju, et al.
Published: (2024)
by: Lin, Jiaju, et al.
Published: (2024)
Similar Items
-
OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models
by: Lin, Gaojie, et al.
Published: (2025) -
AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation
by: Liang, Chao, et al.
Published: (2025) -
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
by: Wang, Zhenzhi, et al.
Published: (2025) -
FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation
by: Zhong, Tianyun, et al.
Published: (2024) -
Semantics-Aware Human Motion Generation from Audio Instructions
by: Wang, Zi-An, et al.
Published: (2025)