CLIPtortionist: Zero-shot Text-driven Deformation for Manufactured 3D Shapes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Xianghao, Sridhar, Srinath, Ritchie, Daniel
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913556621950976
author Xu, Xianghao
Sridhar, Srinath
Ritchie, Daniel
author_facet Xu, Xianghao
Sridhar, Srinath
Ritchie, Daniel
contents We propose a zero-shot text-driven 3D shape deformation system that deforms an input 3D mesh of a manufactured object to fit an input text description. To do this, our system optimizes the parameters of a deformation model to maximize an objective function based on the widely used pre-trained vision language model CLIP. We find that CLIP-based objective functions exhibit many spurious local optima; to circumvent them, we parameterize deformations using a novel deformation model called BoxDefGraph which our system automatically computes from an input mesh, the BoxDefGraph is designed to capture the object aligned rectangular/circular geometry features of most manufactured objects. We then use the CMA-ES global optimization algorithm to maximize our objective, which we find to work better than popular gradient-based optimizers. We demonstrate that our approach produces appealing results and outperforms several baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15199
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CLIPtortionist: Zero-shot Text-driven Deformation for Manufactured 3D Shapes
Xu, Xianghao
Sridhar, Srinath
Ritchie, Daniel
Computer Vision and Pattern Recognition
Graphics
We propose a zero-shot text-driven 3D shape deformation system that deforms an input 3D mesh of a manufactured object to fit an input text description. To do this, our system optimizes the parameters of a deformation model to maximize an objective function based on the widely used pre-trained vision language model CLIP. We find that CLIP-based objective functions exhibit many spurious local optima; to circumvent them, we parameterize deformations using a novel deformation model called BoxDefGraph which our system automatically computes from an input mesh, the BoxDefGraph is designed to capture the object aligned rectangular/circular geometry features of most manufactured objects. We then use the CMA-ES global optimization algorithm to maximize our objective, which we find to work better than popular gradient-based optimizers. We demonstrate that our approach produces appealing results and outperforms several baselines.
title CLIPtortionist: Zero-shot Text-driven Deformation for Manufactured 3D Shapes
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2410.15199