Task-Oriented Data Synthesis and Control-Rectify Sampling for Remote Sensing Semantic Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Yunkai, Zhang, Yudong, Zhang, Kunquan, Zhang, Jinxiao, Chen, Xinying, Fu, Haohuan, Dong, Runmin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915887845474304
author Yang, Yunkai
Zhang, Yudong
Zhang, Kunquan
Zhang, Jinxiao
Chen, Xinying
Fu, Haohuan
Dong, Runmin
author_facet Yang, Yunkai
Zhang, Yudong
Zhang, Kunquan
Zhang, Jinxiao
Chen, Xinying
Fu, Haohuan
Dong, Runmin
contents With the rapid progress of controllable generation, training data synthesis has become a promising way to expand labeled datasets and alleviate manual annotation in remote sensing (RS). However, the complexity of semantic mask control and the uncertainty of sampling quality often limit the utility of synthetic data in downstream semantic segmentation tasks. To address these challenges, we propose a task-oriented data synthesis framework (TODSynth), including a Multimodal Diffusion Transformer (MM-DiT) with unified triple attention and a plug-and-play sampling strategy guided by task feedback. Built upon the powerful DiT-based generative foundation model, we systematically evaluate different control schemes, showing that a text-image-mask joint attention scheme combined with full fine-tuning of the image and mask branches significantly enhances the effectiveness of RS semantic segmentation data synthesis, particularly in few-shot and complex-scene scenarios. Furthermore, we propose a control-rectify flow matching (CRFM) method, which dynamically adjusts sampling directions guided by semantic loss during the early high-plasticity stage, mitigating the instability of generated images and bridging the gap between synthetic data and downstream segmentation tasks. Extensive experiments demonstrate that our approach consistently outperforms state-of-the-art controllable generation methods, producing more stable and task-oriented synthetic data for RS semantic segmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16740
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Task-Oriented Data Synthesis and Control-Rectify Sampling for Remote Sensing Semantic Segmentation
Yang, Yunkai
Zhang, Yudong
Zhang, Kunquan
Zhang, Jinxiao
Chen, Xinying
Fu, Haohuan
Dong, Runmin
Computer Vision and Pattern Recognition
With the rapid progress of controllable generation, training data synthesis has become a promising way to expand labeled datasets and alleviate manual annotation in remote sensing (RS). However, the complexity of semantic mask control and the uncertainty of sampling quality often limit the utility of synthetic data in downstream semantic segmentation tasks. To address these challenges, we propose a task-oriented data synthesis framework (TODSynth), including a Multimodal Diffusion Transformer (MM-DiT) with unified triple attention and a plug-and-play sampling strategy guided by task feedback. Built upon the powerful DiT-based generative foundation model, we systematically evaluate different control schemes, showing that a text-image-mask joint attention scheme combined with full fine-tuning of the image and mask branches significantly enhances the effectiveness of RS semantic segmentation data synthesis, particularly in few-shot and complex-scene scenarios. Furthermore, we propose a control-rectify flow matching (CRFM) method, which dynamically adjusts sampling directions guided by semantic loss during the early high-plasticity stage, mitigating the instability of generated images and bridging the gap between synthetic data and downstream segmentation tasks. Extensive experiments demonstrate that our approach consistently outperforms state-of-the-art controllable generation methods, producing more stable and task-oriented synthetic data for RS semantic segmentation.
title Task-Oriented Data Synthesis and Control-Rectify Sampling for Remote Sensing Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.16740