Learn2Fold: Structured Origami Generation with World Model Planning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Yanjia, Chen, Yunuo, Jiang, Ying, Han, Jinru, Tu, Zhengzhong, Yang, Yin, Jiang, Chenfanfu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908934250430464
author Huang, Yanjia
Chen, Yunuo
Jiang, Ying
Han, Jinru
Tu, Zhengzhong
Yang, Yin
Jiang, Chenfanfu
author_facet Huang, Yanjia
Chen, Yunuo
Jiang, Ying
Han, Jinru
Tu, Zhengzhong
Yang, Yin
Jiang, Chenfanfu
contents The ability to transform a flat sheet into a complex three-dimensional structure is a fundamental test of physical intelligence. Unlike cloth manipulation, origami is governed by strict geometric axioms and hard kinematic constraints, where a single invalid crease or collision can invalidate the entire folding sequence. As a result, origami demands long-horizon constructive reasoning that jointly satisfies precise physical laws and high-level semantic intent. Existing approaches fall into two disjoint paradigms: optimization-based methods enforce physical validity but require dense, precisely specified inputs, making them unsuitable for sparse natural language descriptions, while generative foundation models excel at semantic and perceptual synthesis yet fail to produce long-horizon, physics-consistent folding processes. Consequently, generating valid origami folding sequences directly from text remains an open challenge. To address this gap, we introduce Learn2Fold, a neuro-symbolic framework that formulates origami folding as conditional program induction over a crease-pattern graph. Our key insight is to decouple semantic proposal from physical verification. A large language model generates candidate folding programs from abstract text prompts, while a learned graph-structured world model serves as a differentiable surrogate simulator that predicts physical feasibility and failure modes before execution. Integrated within a lookahead planning loop, Learn2Fold enables robust generation of physically valid folding sequences for complex and out-of-distribution patterns, demonstrating that effective spatial intelligence arises from the synergy between symbolic reasoning and grounded physical simulation.
format Preprint
id arxiv_https___arxiv_org_abs_2603_29585
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learn2Fold: Structured Origami Generation with World Model Planning
Huang, Yanjia
Chen, Yunuo
Jiang, Ying
Han, Jinru
Tu, Zhengzhong
Yang, Yin
Jiang, Chenfanfu
Graphics
Artificial Intelligence
The ability to transform a flat sheet into a complex three-dimensional structure is a fundamental test of physical intelligence. Unlike cloth manipulation, origami is governed by strict geometric axioms and hard kinematic constraints, where a single invalid crease or collision can invalidate the entire folding sequence. As a result, origami demands long-horizon constructive reasoning that jointly satisfies precise physical laws and high-level semantic intent. Existing approaches fall into two disjoint paradigms: optimization-based methods enforce physical validity but require dense, precisely specified inputs, making them unsuitable for sparse natural language descriptions, while generative foundation models excel at semantic and perceptual synthesis yet fail to produce long-horizon, physics-consistent folding processes. Consequently, generating valid origami folding sequences directly from text remains an open challenge. To address this gap, we introduce Learn2Fold, a neuro-symbolic framework that formulates origami folding as conditional program induction over a crease-pattern graph. Our key insight is to decouple semantic proposal from physical verification. A large language model generates candidate folding programs from abstract text prompts, while a learned graph-structured world model serves as a differentiable surrogate simulator that predicts physical feasibility and failure modes before execution. Integrated within a lookahead planning loop, Learn2Fold enables robust generation of physically valid folding sequences for complex and out-of-distribution patterns, demonstrating that effective spatial intelligence arises from the synergy between symbolic reasoning and grounded physical simulation.
title Learn2Fold: Structured Origami Generation with World Model Planning
topic Graphics
Artificial Intelligence
url https://arxiv.org/abs/2603.29585