OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Jianwen, Zeng, Weihong, Zheng, Zerong, Yang, Jiaqi, Liang, Chao, Liao, Wang, Liang, Han, Zhang, Yuan, Gao, Mingyuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models
by: Lin, Gaojie, et al.
Published: (2025)
by: Lin, Gaojie, et al.
Published: (2025)
AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation
by: Liang, Chao, et al.
Published: (2025)
by: Liang, Chao, et al.
Published: (2025)
InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions
by: Wang, Zhenzhi, et al.
Published: (2025)
by: Wang, Zhenzhi, et al.
Published: (2025)
OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation
by: Zhu, Lei, et al.
Published: (2026)
by: Zhu, Lei, et al.
Published: (2026)
Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency
by: Jiang, Jianwen, et al.
Published: (2024)
by: Jiang, Jianwen, et al.
Published: (2024)
CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention
by: Lin, Gaojie, et al.
Published: (2024)
by: Lin, Gaojie, et al.
Published: (2024)
MeshAvatar: Learning High-quality Triangular Human Avatars from Multi-view Videos
by: Chen, Yushuo, et al.
Published: (2024)
by: Chen, Yushuo, et al.
Published: (2024)
DreamVVT: Mastering Realistic Video Virtual Try-On in the Wild via a Stage-Wise Diffusion Transformer Framework
by: Zuo, Tongchun, et al.
Published: (2025)
by: Zuo, Tongchun, et al.
Published: (2025)
FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation
by: Zhong, Tianyun, et al.
Published: (2024)
by: Zhong, Tianyun, et al.
Published: (2024)
MobilePortrait: Real-Time One-Shot Neural Head Avatars on Mobile Devices
by: Jiang, Jianwen, et al.
Published: (2024)
by: Jiang, Jianwen, et al.
Published: (2024)
Animatable and Relightable Gaussians for High-fidelity Human Avatar Modeling
by: Li, Zhe, et al.
Published: (2023)
by: Li, Zhe, et al.
Published: (2023)
ThinkTuning: Instilling Cognitive Reflections without Distillation
by: RRV, Aswin, et al.
Published: (2025)
by: RRV, Aswin, et al.
Published: (2025)
Instilling Multi-round Thinking to Text-guided Image Generation
by: Zeng, Lidong, et al.
Published: (2024)
by: Zeng, Lidong, et al.
Published: (2024)
The Tonogenesis Continuum in Tibetan: A Computational Investigation
by: Liang, Siyu, et al.
Published: (2025)
by: Liang, Siyu, et al.
Published: (2025)
Instilling Organisational Values in Firefighters through Simulation-Based Training
by: Osman, Nardine, et al.
Published: (2025)
by: Osman, Nardine, et al.
Published: (2025)
DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers
by: Wang, Lizhen, et al.
Published: (2025)
by: Wang, Lizhen, et al.
Published: (2025)
Stability of AI Governance Systems: A Coupled Dynamics Model of Public Trust and Social Disruptions
by: Lai, Jiaqi, et al.
Published: (2026)
by: Lai, Jiaqi, et al.
Published: (2026)
LayGA: Layered Gaussian Avatars for Animatable Clothing Transfer
by: Lin, Siyou, et al.
Published: (2024)
by: Lin, Siyou, et al.
Published: (2024)
Instilling Inductive Biases with Subnetworks
by: Zhang, Enyan, et al.
Published: (2023)
by: Zhang, Enyan, et al.
Published: (2023)
FlowAct-R1: Towards Interactive Humanoid Video Generation
by: Wang, Lizhen, et al.
Published: (2026)
by: Wang, Lizhen, et al.
Published: (2026)
MirrorMind: Empowering OmniScientist with the Expert Perspectives and Collective Knowledge of Human Scientists
by: Zeng, Qingbin, et al.
Published: (2025)
by: Zeng, Qingbin, et al.
Published: (2025)
Superior and Pragmatic Talking Face Generation with Teacher-Student Framework
by: Liang, Chao, et al.
Published: (2024)
by: Liang, Chao, et al.
Published: (2024)
Chronocept: Instilling a Sense of Time in Machines
by: Goel, Krish, et al.
Published: (2025)
by: Goel, Krish, et al.
Published: (2025)
Synergy: End-to-end Concept Model
by: Zheng, Keli, et al.
Published: (2025)
by: Zheng, Keli, et al.
Published: (2025)
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
by: Gan, Qijun, et al.
Published: (2025)
by: Gan, Qijun, et al.
Published: (2025)
Baichuan-Omni-1.5 Technical Report
by: Li, Yadong, et al.
Published: (2025)
by: Li, Yadong, et al.
Published: (2025)
Omni-Judge: Can Omni-LLMs Serve as Human-Aligned Judges for Text-Conditioned Audio-Video Generation?
by: Liang, Susan, et al.
Published: (2026)
by: Liang, Susan, et al.
Published: (2026)
Can Large Language Models Simulate Human Cognition Beyond Behavioral Imitation?
by: Gu, Yuxuan, et al.
Published: (2026)
by: Gu, Yuxuan, et al.
Published: (2026)
InsertNeRF: Instilling Generalizability into NeRF with HyperNet Modules
by: Bao, Yanqi, et al.
Published: (2023)
by: Bao, Yanqi, et al.
Published: (2023)
VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
by: Yang, Baoyao, et al.
Published: (2025)
by: Yang, Baoyao, et al.
Published: (2025)
FLINGO -- Instilling ASP Expressiveness into Linear Integer Constraints
by: Fandinno, Jorge, et al.
Published: (2026)
by: Fandinno, Jorge, et al.
Published: (2026)
ToolDreamer: Instilling LLM Reasoning Into Tool Retrievers
by: Sengupta, Saptarshi, et al.
Published: (2025)
by: Sengupta, Saptarshi, et al.
Published: (2025)
UltrAvatar: A Realistic Animatable 3D Avatar Diffusion Model with Authenticity Guided Textures
by: Zhou, Mingyuan, et al.
Published: (2024)
by: Zhou, Mingyuan, et al.
Published: (2024)
Accelerating Discovery of Polyimides with Intrinsic Microporosity for Membrane‐Based Gas Separation: Synergizing Physics‐Informed Performance Metrics and Active Learning
by: Mao Wang, et al.
Published: (2024)
by: Mao Wang, et al.
Published: (2024)
Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
by: Shao, Ruizhi, et al.
Published: (2024)
by: Shao, Ruizhi, et al.
Published: (2024)
AvatarGO: Zero-shot 4D Human-Object Interaction Generation and Animation
by: Cao, Yukang, et al.
Published: (2024)
by: Cao, Yukang, et al.
Published: (2024)
PAGED: A Benchmark for Procedural Graphs Extraction from Documents
by: Du, Weihong, et al.
Published: (2024)
by: Du, Weihong, et al.
Published: (2024)
Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars
by: Zhang, Youliang, et al.
Published: (2026)
by: Zhang, Youliang, et al.
Published: (2026)
SimAvatar: Simulation-Ready Avatars with Layered Hair and Clothing
by: Li, Xueting, et al.
Published: (2024)
by: Li, Xueting, et al.
Published: (2024)
Continuum Mechanics Modeling of Flexible Spring Joints in Surgical Robots
by: Botian Sun, et al.
Published: (2025)
by: Botian Sun, et al.
Published: (2025)
Similar Items
-
OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models
by: Lin, Gaojie, et al.
Published: (2025) -
AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation
by: Liang, Chao, et al.
Published: (2025) -
InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions
by: Wang, Zhenzhi, et al.
Published: (2025) -
OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation
by: Zhu, Lei, et al.
Published: (2026) -
Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency
by: Jiang, Jianwen, et al.
Published: (2024)