Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Kunyu, Ma, Yue, Zhang, Xinhua, Liu, Boshi, Yuluo, Yikuang, Zhang, Yinhan, Liu, Runtao, Liu, Hongyu, Qin, Zhiyuan, Mo, Shanhui, Chen, Qifeng, Wang, Zeyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Follow-Your-Creation: Empowering 4D Creation through Video Inpainting
by: Ma, Yue, et al.
Published: (2025)
by: Ma, Yue, et al.
Published: (2025)
Follow-Your-Color: Multi-Instance Sketch Colorization
by: Zhang, Yinhan, et al.
Published: (2025)
by: Zhang, Yinhan, et al.
Published: (2025)
Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region Control
by: Long, Zeqian, et al.
Published: (2025)
by: Long, Zeqian, et al.
Published: (2025)
EVCtrl: Efficient Control Adapter for Visual Generation
by: Yang, Zixiang, et al.
Published: (2025)
by: Yang, Zixiang, et al.
Published: (2025)
Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning
by: Ma, Yue, et al.
Published: (2025)
by: Ma, Yue, et al.
Published: (2025)
LongVideoAgent: Multi-Agent Reasoning with Long Videos
by: Liu, Runtao, et al.
Published: (2025)
by: Liu, Runtao, et al.
Published: (2025)
InstanceAnimator: Multi-Instance Sketch Video Colorization
by: Zhang, Yinhan, et al.
Published: (2026)
by: Zhang, Yinhan, et al.
Published: (2026)
EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation
by: Ma, Yue, et al.
Published: (2026)
by: Ma, Yue, et al.
Published: (2026)
GR-Gaussian: Graph-Based Radiative Gaussian Splatting for Sparse-View CT Reconstruction
by: Yuluo, Yikuang, et al.
Published: (2025)
by: Yuluo, Yikuang, et al.
Published: (2025)
Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation
by: Ma, Yue, et al.
Published: (2024)
by: Ma, Yue, et al.
Published: (2024)
Follow-Your-Emoji-Faster: Towards Efficient, Fine-Controllable, and Expressive Freestyle Portrait Animation
by: Ma, Yue, et al.
Published: (2025)
by: Ma, Yue, et al.
Published: (2025)
AvatarPointillist: AutoRegressive 4D Gaussian Avatarization
by: Liu, Hongyu, et al.
Published: (2026)
by: Liu, Hongyu, et al.
Published: (2026)
Effective Training Data Synthesis for Improving MLLM Chart Understanding
by: Yang, Yuwei, et al.
Published: (2025)
by: Yang, Yuwei, et al.
Published: (2025)
OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents
by: Wang, Zihao, et al.
Published: (2024)
by: Wang, Zihao, et al.
Published: (2024)
OPEx: A Component-Wise Analysis of LLM-Centric Agents in Embodied Instruction Following
by: Shi, Haochen, et al.
Published: (2024)
by: Shi, Haochen, et al.
Published: (2024)
Neurobiological Synergy of Plant and Animal Sources of Omega‐3 and Exercise in Aging: Implications for Molecular Signaling, Memory, Spatial Learning, and Brain Function
by: Yikuang Hao, et al.
Published: (2025)
by: Yikuang Hao, et al.
Published: (2025)
EditWorld: Simulating World Dynamics for Instruction-Following Image Editing
by: Yang, Ling, et al.
Published: (2024)
by: Yang, Ling, et al.
Published: (2024)
MR-MLLM: Mutual Reinforcement of Multimodal Comprehension and Vision Perception
by: Wang, Guanqun, et al.
Published: (2024)
by: Wang, Guanqun, et al.
Published: (2024)
Control and Realism: Best of Both Worlds in Layout-to-Image without Training
by: Li, Bonan, et al.
Published: (2025)
by: Li, Bonan, et al.
Published: (2025)
TIIF-Bench: How Does Your T2I Model Follow Your Instructions?
by: Wei, Xinyu, et al.
Published: (2025)
by: Wei, Xinyu, et al.
Published: (2025)
Follow-Your-Canvas: Higher-Resolution Video Outpainting with Extensive Content Generation
by: Chen, Qihua, et al.
Published: (2024)
by: Chen, Qihua, et al.
Published: (2024)
Fake it till You Make it: Reward Modeling as Discriminative Prediction
by: Liu, Runtao, et al.
Published: (2025)
by: Liu, Runtao, et al.
Published: (2025)
Latent Guard: a Safety Framework for Text-to-image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
SOTOPIA-$Ω$: Dynamic Strategy Injection Learning and Social Instruction Following Evaluation for Social Agents
by: Zhang, Wenyuan, et al.
Published: (2025)
by: Zhang, Wenyuan, et al.
Published: (2025)
CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation
by: Ma, Xinbei, et al.
Published: (2024)
by: Ma, Xinbei, et al.
Published: (2024)
Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw
by: Wang, Zijun, et al.
Published: (2026)
by: Wang, Zijun, et al.
Published: (2026)
RenderFlow: Single-Step Neural Rendering via Flow Matching
by: Zhang, Shenghao, et al.
Published: (2026)
by: Zhang, Shenghao, et al.
Published: (2026)
Artic: AI-oriented Real-time Communication for MLLM Video Assistant
by: Wu, Jiangkai, et al.
Published: (2026)
by: Wu, Jiangkai, et al.
Published: (2026)
DiT4Edit: Diffusion Transformer for Image Editing
by: Feng, Kunyu, et al.
Published: (2024)
by: Feng, Kunyu, et al.
Published: (2024)
OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents
by: Cheng, Pengzhou, et al.
Published: (2025)
by: Cheng, Pengzhou, et al.
Published: (2025)
LLMs Meet Multimodal Generation and Editing: A Survey
by: He, Yingqing, et al.
Published: (2024)
by: He, Yingqing, et al.
Published: (2024)
Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts
by: Ma, Yue, et al.
Published: (2024)
by: Ma, Yue, et al.
Published: (2024)
GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents
by: Cai, Shaofei, et al.
Published: (2024)
by: Cai, Shaofei, et al.
Published: (2024)
HiPrompt: Tuning-free Higher-Resolution Generation with Hierarchical MLLM Prompts
by: Liu, Xinyu, et al.
Published: (2024)
by: Liu, Xinyu, et al.
Published: (2024)
PanicFI: An Infrastructure for Fixing Panic Bugs in Real-World Rust Programs
by: Ni, Yunbo, et al.
Published: (2024)
by: Ni, Yunbo, et al.
Published: (2024)
MultiMotion: Multi Subject Video Motion Transfer via Video Diffusion Transformer
by: Liu, Penghui, et al.
Published: (2025)
by: Liu, Penghui, et al.
Published: (2025)
MiniCPM-V: A GPT-4V Level MLLM on Your Phone
by: Yao, Yuan, et al.
Published: (2024)
by: Yao, Yuan, et al.
Published: (2024)
VideoDPO: Omni-Preference Alignment for Video Diffusion Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
Dermatoscopic features of vulvar lichen sclerosus in children: A retrospective study
by: Yuyang Han, et al.
Published: (2024)
by: Yuyang Han, et al.
Published: (2024)
FollowTable: A Benchmark for Instruction-Following Table Retrieval
by: Jin, Rihui, et al.
Published: (2026)
by: Jin, Rihui, et al.
Published: (2026)
Similar Items
-
Follow-Your-Creation: Empowering 4D Creation through Video Inpainting
by: Ma, Yue, et al.
Published: (2025) -
Follow-Your-Color: Multi-Instance Sketch Colorization
by: Zhang, Yinhan, et al.
Published: (2025) -
Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region Control
by: Long, Zeqian, et al.
Published: (2025) -
EVCtrl: Efficient Control Adapter for Visual Generation
by: Yang, Zixiang, et al.
Published: (2025) -
Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning
by: Ma, Yue, et al.
Published: (2025)