CAD-Coder: Text-to-CAD Generation with Chain-of-Thought and Geometric Reward

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Guan, Yandong, Wang, Xilin, Xing, Ximing, Zhang, Jing, Xu, Dong, Yu, Qian
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910219738546176
author Guan, Yandong
Wang, Xilin
Xing, Ximing
Zhang, Jing
Xu, Dong
Yu, Qian
author_facet Guan, Yandong
Wang, Xilin
Xing, Ximing
Zhang, Jing
Xu, Dong
Yu, Qian
contents In this work, we introduce CAD-Coder, a novel framework that reformulates text-to-CAD as the generation of CadQuery scripts - a Python-based, parametric CAD language. This representation enables direct geometric validation, a richer modeling vocabulary, and seamless integration with existing LLMs. To further enhance code validity and geometric fidelity, we propose a two-stage learning pipeline: (1) supervised fine-tuning on paired text-CadQuery data, and (2) reinforcement learning with Group Reward Policy Optimization (GRPO), guided by a CAD-specific reward comprising both a geometric reward (Chamfer Distance) and a format reward. We also introduce a chain-of-thought (CoT) planning process to improve model reasoning, and construct a large-scale, high-quality dataset of 110K text-CadQuery-3D model triplets and 1.5K CoT samples via an automated pipeline. Extensive experiments demonstrate that CAD-Coder enables LLMs to generate diverse, valid, and complex CAD models directly from natural language, advancing the state of the art of text-to-CAD generation and geometric reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19713
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CAD-Coder: Text-to-CAD Generation with Chain-of-Thought and Geometric Reward
Guan, Yandong
Wang, Xilin
Xing, Ximing
Zhang, Jing
Xu, Dong
Yu, Qian
Graphics
In this work, we introduce CAD-Coder, a novel framework that reformulates text-to-CAD as the generation of CadQuery scripts - a Python-based, parametric CAD language. This representation enables direct geometric validation, a richer modeling vocabulary, and seamless integration with existing LLMs. To further enhance code validity and geometric fidelity, we propose a two-stage learning pipeline: (1) supervised fine-tuning on paired text-CadQuery data, and (2) reinforcement learning with Group Reward Policy Optimization (GRPO), guided by a CAD-specific reward comprising both a geometric reward (Chamfer Distance) and a format reward. We also introduce a chain-of-thought (CoT) planning process to improve model reasoning, and construct a large-scale, high-quality dataset of 110K text-CadQuery-3D model triplets and 1.5K CoT samples via an automated pipeline. Extensive experiments demonstrate that CAD-Coder enables LLMs to generate diverse, valid, and complex CAD models directly from natural language, advancing the state of the art of text-to-CAD generation and geometric reasoning.
title CAD-Coder: Text-to-CAD Generation with Chain-of-Thought and Geometric Reward
topic Graphics
url https://arxiv.org/abs/2505.19713