Saved in:
Bibliographic Details
Main Author: Guo, Qiushi
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2407.08151
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915295281545216
author Guo, Qiushi
author_facet Guo, Qiushi
contents Data augmentation remains a widely utilized technique in deep learning, particularly in tasks such as image classification, semantic segmentation, and object detection. Among them, Copy-Paste is a simple yet effective method and gain great attention recently. However, existing Copy-Paste often overlook contextual relevance between source and target images, resulting in inconsistencies in generated outputs. To address this challenge, we propose a context-aware approach that integrates Bidirectional Latent Information Propagation (BLIP) for content extraction from source images. By matching extracted content information with category information, our method ensures cohesive integration of target objects using Segment Anything Model (SAM) and You Only Look Once (YOLO). This approach eliminates the need for manual annotation, offering an automated and user-friendly solution. Experimental evaluations across diverse datasets demonstrate the effectiveness of our method in enhancing data diversity and generating high-quality pseudo-images across various computer vision tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2407_08151
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enrich the content of the image Using Context-Aware Copy Paste
Guo, Qiushi
Computer Vision and Pattern Recognition
Data augmentation remains a widely utilized technique in deep learning, particularly in tasks such as image classification, semantic segmentation, and object detection. Among them, Copy-Paste is a simple yet effective method and gain great attention recently. However, existing Copy-Paste often overlook contextual relevance between source and target images, resulting in inconsistencies in generated outputs. To address this challenge, we propose a context-aware approach that integrates Bidirectional Latent Information Propagation (BLIP) for content extraction from source images. By matching extracted content information with category information, our method ensures cohesive integration of target objects using Segment Anything Model (SAM) and You Only Look Once (YOLO). This approach eliminates the need for manual annotation, offering an automated and user-friendly solution. Experimental evaluations across diverse datasets demonstrate the effectiveness of our method in enhancing data diversity and generating high-quality pseudo-images across various computer vision tasks.
title Enrich the content of the image Using Context-Aware Copy Paste
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.08151