OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions
Fuente:
arXiv
Saved in:
| Main Authors: | Bu, Wendong, Pan, Kaihang, Lin, Yuze, Li, Jiacheng, Shen, Kai, Zhang, Wenqiao, Li, Juncheng, Xiao, Jun, Tang, Siliang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness
by: Qiu, Haiyi, et al.
Published: (2026)
by: Qiu, Haiyi, et al.
Published: (2026)
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL
by: Pan, Kaihang, et al.
Published: (2025)
by: Pan, Kaihang, et al.
Published: (2025)
Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning
by: Pan, Kaihang, et al.
Published: (2025)
by: Pan, Kaihang, et al.
Published: (2025)
What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
by: Bu, Wendong, et al.
Published: (2025)
by: Bu, Wendong, et al.
Published: (2025)
WiseEdit: Benchmarking Cognition- and Creativity-Informed Image Editing
by: Pan, Kaihang, et al.
Published: (2025)
by: Pan, Kaihang, et al.
Published: (2025)
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
by: Li, Juncheng, et al.
Published: (2023)
by: Li, Juncheng, et al.
Published: (2023)
CORE: Code-based Inverse Self-Training Framework with Graph Expansion for Virtual Agents
by: Wang, Keyu, et al.
Published: (2026)
by: Wang, Keyu, et al.
Published: (2026)
X-MoGen: Unified Motion Generation across Humans and Animals
by: Wang, Xuan, et al.
Published: (2025)
by: Wang, Xuan, et al.
Published: (2025)
InstructSAM: Segment Any Instance with Any Instructions
by: Yuan, Yuqian, et al.
Published: (2026)
by: Yuan, Yuqian, et al.
Published: (2026)
Learning to Adapt: Self-Improving Web Agent via Cognitive-Aware Exploration
by: Chen, Weile, et al.
Published: (2026)
by: Chen, Weile, et al.
Published: (2026)
SnapMoGen: Human Motion Generation from Expressive Texts
by: Guo, Chuan, et al.
Published: (2025)
by: Guo, Chuan, et al.
Published: (2025)
Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark
by: Miao, Bingchen, et al.
Published: (2025)
by: Miao, Bingchen, et al.
Published: (2025)
Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration
by: Pan, Kaihang, et al.
Published: (2024)
by: Pan, Kaihang, et al.
Published: (2024)
I3: Intent-Introspective Retrieval Conditioned on Instructions
by: Pan, Kaihang, et al.
Published: (2023)
by: Pan, Kaihang, et al.
Published: (2023)
Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining
by: Ge, Zhiqi, et al.
Published: (2024)
by: Ge, Zhiqi, et al.
Published: (2024)
Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
by: Fan, Zhaoyu, et al.
Published: (2025)
by: Fan, Zhaoyu, et al.
Published: (2025)
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
by: Li, Yuan-Ming, et al.
Published: (2025)
by: Li, Yuan-Ming, et al.
Published: (2025)
Align$^2$LLaVA: Cascaded Human and Large Language Model Preference Alignment for Multi-modal Instruction Curation
by: Huang, Hongzhe, et al.
Published: (2024)
by: Huang, Hongzhe, et al.
Published: (2024)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
by: Chow, Wei, et al.
Published: (2024)
by: Chow, Wei, et al.
Published: (2024)
AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
by: Yu, Qifan, et al.
Published: (2024)
by: Yu, Qifan, et al.
Published: (2024)
OmniMotion: Multimodal Motion Generation with Continuous Masked Autoregression
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
OmniGen: Unified Image Generation
by: Xiao, Shitao, et al.
Published: (2024)
by: Xiao, Shitao, et al.
Published: (2024)
CoMo: Compositional Motion Customization for Text-to-Video Generation
by: Xu, Youcan, et al.
Published: (2025)
by: Xu, Youcan, et al.
Published: (2025)
SimMotionEdit: Text-Based Human Motion Editing with Motion Similarity Prediction
by: Li, Zhengyuan, et al.
Published: (2025)
by: Li, Zhengyuan, et al.
Published: (2025)
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
by: Miao, Bingchen, et al.
Published: (2024)
by: Miao, Bingchen, et al.
Published: (2024)
MoRL: Reinforced Reasoning for Unified Motion Understanding and Generation
by: Wang, Hongpeng, et al.
Published: (2026)
by: Wang, Hongpeng, et al.
Published: (2026)
Bridging Local Details and Global Context in Text-Attributed Graphs
by: Wang, Yaoke, et al.
Published: (2024)
by: Wang, Yaoke, et al.
Published: (2024)
OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning
by: Pan, Kaihang, et al.
Published: (2026)
by: Pan, Kaihang, et al.
Published: (2026)
Multi-Scale Incremental Modeling for Enhanced Human Motion Prediction in Human-Robot Collaboration
by: Zou, Juncheng
Published: (2024)
by: Zou, Juncheng
Published: (2024)
LaMoGen: Laban Movement-Guided Diffusion for Text-to-Motion Generation
by: Kim, Heechang, et al.
Published: (2025)
by: Kim, Heechang, et al.
Published: (2025)
CrowdMoGen: Zero-Shot Text-Driven Collective Motion Generation
by: Cao, Yukang, et al.
Published: (2024)
by: Cao, Yukang, et al.
Published: (2024)
UniMuMo: Unified Text, Music and Motion Generation
by: Yang, Han, et al.
Published: (2024)
by: Yang, Han, et al.
Published: (2024)
OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
by: Li, Qingyun, et al.
Published: (2024)
by: Li, Qingyun, et al.
Published: (2024)
MoA: Heterogeneous Mixture of Adapters for Parameter-Efficient Fine-Tuning of Large Language Models
by: Cao, Jie, et al.
Published: (2025)
by: Cao, Jie, et al.
Published: (2025)
UniMoGen: Universal Motion Generation
by: Khani, Aliasghar, et al.
Published: (2025)
by: Khani, Aliasghar, et al.
Published: (2025)
LASER: Tuning-Free LLM-Driven Attention Control for Efficient Text-conditioned Image-to-Animation
by: Zheng, Haoyu, et al.
Published: (2024)
by: Zheng, Haoyu, et al.
Published: (2024)
Human Motion Instruction Tuning
by: Li, Lei, et al.
Published: (2024)
by: Li, Lei, et al.
Published: (2024)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
by: Qin, Bosheng, et al.
Published: (2023)
by: Qin, Bosheng, et al.
Published: (2023)
FreeMotion: A Unified Framework for Number-free Text-to-Motion Synthesis
by: Fan, Ke, et al.
Published: (2024)
by: Fan, Ke, et al.
Published: (2024)
Infinite Motion: Extended Motion Generation via Long Text Instructions
by: Li, Mengtian, et al.
Published: (2024)
by: Li, Mengtian, et al.
Published: (2024)
Similar Items
-
SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness
by: Qiu, Haiyi, et al.
Published: (2026) -
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL
by: Pan, Kaihang, et al.
Published: (2025) -
Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning
by: Pan, Kaihang, et al.
Published: (2025) -
What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
by: Bu, Wendong, et al.
Published: (2025) -
WiseEdit: Benchmarking Cognition- and Creativity-Informed Image Editing
by: Pan, Kaihang, et al.
Published: (2025)