ClothTransformer: Unified Latent-Space Transformers for Scalable Cloth Simulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yu, Shao, Yidi, Ouyang, Wenqi, Lan, Yushi, Liang, Zhexin, Wu, Chengrui, Xu, Xudong, Pan, Xingang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913166426898432
author Zhang, Yu
Shao, Yidi
Ouyang, Wenqi
Lan, Yushi
Liang, Zhexin
Wu, Chengrui
Xu, Xudong
Pan, Xingang
author_facet Zhang, Yu
Shao, Yidi
Ouyang, Wenqi
Lan, Yushi
Liang, Zhexin
Wu, Chengrui
Xu, Xudong
Pan, Xingang
contents Unified and scalable Transformers have recently achieved remarkable success in modeling diverse phenomena traditionally associated with computer graphics, such as 3D visual effects, rendering processes, and motion in videos. In this work, we take a step further by investigating whether modern Transformer techniques can tackle the challenging task of cloth simulation. To this end, we present ClothTransformer, a framework that reformulates cloth simulation as autoregressive sequence modeling in a learned latent space. Existing neural cloth simulators are largely specialized to single scenarios, intrinsically coupled to the mesh discretization, and lack robust collision handling. Our approach addresses these limitations through three contributions: (1) a unified Transformer architecture that handles diverse scenarios -- body-driven garments, robotic manipulation, and free-fall collisions -- under a single model and achieves approximately $4$--$9{\times}$ lower error than prior state-of-the-art methods across all scenarios; (2) a scalable latent-space formulation that compresses arbitrary-resolution meshes into a fixed-size set of latent tokens, making temporal dynamics computation independent of mesh resolution; and (3) a diverse-scenario high-fidelity penetration-free dataset of ${\sim}$493.4k frames spanning all three settings, which enables a differentiable Continuous Collision Detection (CCD) module to suppress penetration artifacts.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27852
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ClothTransformer: Unified Latent-Space Transformers for Scalable Cloth Simulation
Zhang, Yu
Shao, Yidi
Ouyang, Wenqi
Lan, Yushi
Liang, Zhexin
Wu, Chengrui
Xu, Xudong
Pan, Xingang
Graphics
Computer Vision and Pattern Recognition
Unified and scalable Transformers have recently achieved remarkable success in modeling diverse phenomena traditionally associated with computer graphics, such as 3D visual effects, rendering processes, and motion in videos. In this work, we take a step further by investigating whether modern Transformer techniques can tackle the challenging task of cloth simulation. To this end, we present ClothTransformer, a framework that reformulates cloth simulation as autoregressive sequence modeling in a learned latent space. Existing neural cloth simulators are largely specialized to single scenarios, intrinsically coupled to the mesh discretization, and lack robust collision handling. Our approach addresses these limitations through three contributions: (1) a unified Transformer architecture that handles diverse scenarios -- body-driven garments, robotic manipulation, and free-fall collisions -- under a single model and achieves approximately $4$--$9{\times}$ lower error than prior state-of-the-art methods across all scenarios; (2) a scalable latent-space formulation that compresses arbitrary-resolution meshes into a fixed-size set of latent tokens, making temporal dynamics computation independent of mesh resolution; and (3) a diverse-scenario high-fidelity penetration-free dataset of ${\sim}$493.4k frames spanning all three settings, which enables a differentiable Continuous Collision Detection (CCD) module to suppress penetration artifacts.
title ClothTransformer: Unified Latent-Space Transformers for Scalable Cloth Simulation
topic Graphics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.27852