Teaching an Agent to Sketch One Part at a Time
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866910160816963584 |
|---|---|
| author | Du, Xiaodan Xu, Ruize Yunis, David Vinker, Yael Shakhnarovich, Greg |
| author_facet | Du, Xiaodan Xu, Ruize Yunis, David Vinker, Yael Shakhnarovich, Greg |
| contents | We develop a method for producing vector sketches one part at a time. To do this, we train a multi-modal language model-based agent using a novel multi-turn process-reward reinforcement learning following supervised fine-tuning. Our approach is enabled by a new dataset we call ControlSketch-Part, containing rich part-level annotations for sketches, obtained using a novel, generic automatic annotation pipeline that segments vector sketches into semantic parts and assigns paths to parts with a structured multi-stage labeling process. Our results indicate that incorporating structured part-level data and providing agent with the visual feedback through the process enables interpretable, controllable, and locally editable text-to-vector sketch generation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_19500 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Teaching an Agent to Sketch One Part at a Time Du, Xiaodan Xu, Ruize Yunis, David Vinker, Yael Shakhnarovich, Greg Artificial Intelligence Computer Vision and Pattern Recognition Graphics Machine Learning We develop a method for producing vector sketches one part at a time. To do this, we train a multi-modal language model-based agent using a novel multi-turn process-reward reinforcement learning following supervised fine-tuning. Our approach is enabled by a new dataset we call ControlSketch-Part, containing rich part-level annotations for sketches, obtained using a novel, generic automatic annotation pipeline that segments vector sketches into semantic parts and assigns paths to parts with a structured multi-stage labeling process. Our results indicate that incorporating structured part-level data and providing agent with the visual feedback through the process enables interpretable, controllable, and locally editable text-to-vector sketch generation. |
| title | Teaching an Agent to Sketch One Part at a Time |
| topic | Artificial Intelligence Computer Vision and Pattern Recognition Graphics Machine Learning |
| url | https://arxiv.org/abs/2603.19500 |