MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Yiren, Liu, Cheng, Shou, Mike Zheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917913252855808
author Song, Yiren
Liu, Cheng
Shou, Mike Zheng
author_facet Song, Yiren
Liu, Cheng
Shou, Mike Zheng
contents A hallmark of human intelligence is the ability to create complex artifacts through structured multi-step processes. Generating procedural tutorials with AI is a longstanding but challenging goal, facing three key obstacles: (1) scarcity of multi-task procedural datasets, (2) maintaining logical continuity and visual consistency between steps, and (3) generalizing across multiple domains. To address these challenges, we propose a multi-domain dataset covering 21 tasks with over 24,000 procedural sequences. Building upon this foundation, we introduce MakeAnything, a framework based on the diffusion transformer (DIT), which leverages fine-tuning to activate the in-context capabilities of DIT for generating consistent procedural sequences. We introduce asymmetric low-rank adaptation (LoRA) for image generation, which balances generalization capabilities and task-specific performance by freezing encoder parameters while adaptively tuning decoder layers. Additionally, our ReCraft model enables image-to-process generation through spatiotemporal consistency constraints, allowing static images to be decomposed into plausible creation sequences. Extensive experiments demonstrate that MakeAnything surpasses existing methods, setting new performance benchmarks for procedural generation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2502_01572
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation
Song, Yiren
Liu, Cheng
Shou, Mike Zheng
Computer Vision and Pattern Recognition
A hallmark of human intelligence is the ability to create complex artifacts through structured multi-step processes. Generating procedural tutorials with AI is a longstanding but challenging goal, facing three key obstacles: (1) scarcity of multi-task procedural datasets, (2) maintaining logical continuity and visual consistency between steps, and (3) generalizing across multiple domains. To address these challenges, we propose a multi-domain dataset covering 21 tasks with over 24,000 procedural sequences. Building upon this foundation, we introduce MakeAnything, a framework based on the diffusion transformer (DIT), which leverages fine-tuning to activate the in-context capabilities of DIT for generating consistent procedural sequences. We introduce asymmetric low-rank adaptation (LoRA) for image generation, which balances generalization capabilities and task-specific performance by freezing encoder parameters while adaptively tuning decoder layers. Additionally, our ReCraft model enables image-to-process generation through spatiotemporal consistency constraints, allowing static images to be decomposed into plausible creation sequences. Extensive experiments demonstrate that MakeAnything surpasses existing methods, setting new performance benchmarks for procedural generation tasks.
title MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.01572