ScribbleSense: Generative Scribble-Based Texture Editing with Intent Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yudi, Geng, Yeming, Zhang, Lei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918314518773760
author Zhang, Yudi
Geng, Yeming
Zhang, Lei
author_facet Zhang, Yudi
Geng, Yeming
Zhang, Lei
contents Interactive 3D model texture editing presents enhanced opportunities for creating 3D assets, with freehand drawing style offering the most intuitive experience. However, existing methods primarily support sketch-based interactions for outlining, while the utilization of coarse-grained scribble-based interaction remains limited. Furthermore, current methodologies often encounter challenges due to the abstract nature of scribble instructions, which can result in ambiguous editing intentions and unclear target semantic locations. To address these issues, we propose ScribbleSense, an editing method that combines multimodal large language models (MLLMs) and image generation models to effectively resolve these challenges. We leverage the visual capabilities of MLLMs to predict the editing intent behind the scribbles. Once the semantic intent of the scribble is discerned, we employ globally generated images to extract local texture details, thereby anchoring local semantics and alleviating ambiguities concerning the target semantic locations. Experimental results indicate that our method effectively leverages the strengths of MLLMs, achieving state-of-the-art interactive editing performance for scribble-based texture editing.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22455
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ScribbleSense: Generative Scribble-Based Texture Editing with Intent Prediction
Zhang, Yudi
Geng, Yeming
Zhang, Lei
Computer Vision and Pattern Recognition
Interactive 3D model texture editing presents enhanced opportunities for creating 3D assets, with freehand drawing style offering the most intuitive experience. However, existing methods primarily support sketch-based interactions for outlining, while the utilization of coarse-grained scribble-based interaction remains limited. Furthermore, current methodologies often encounter challenges due to the abstract nature of scribble instructions, which can result in ambiguous editing intentions and unclear target semantic locations. To address these issues, we propose ScribbleSense, an editing method that combines multimodal large language models (MLLMs) and image generation models to effectively resolve these challenges. We leverage the visual capabilities of MLLMs to predict the editing intent behind the scribbles. Once the semantic intent of the scribble is discerned, we employ globally generated images to extract local texture details, thereby anchoring local semantics and alleviating ambiguities concerning the target semantic locations. Experimental results indicate that our method effectively leverages the strengths of MLLMs, achieving state-of-the-art interactive editing performance for scribble-based texture editing.
title ScribbleSense: Generative Scribble-Based Texture Editing with Intent Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.22455