Unifying Image Processing as Visual Prompting Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yihao, Chen, Xiangyu, Ma, Xianzheng, Wang, Xintao, Zhou, Jiantao, Qiao, Yu, Dong, Chao
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914687371706368
author Liu, Yihao
Chen, Xiangyu
Ma, Xianzheng
Wang, Xintao
Zhou, Jiantao
Qiao, Yu
Dong, Chao
author_facet Liu, Yihao
Chen, Xiangyu
Ma, Xianzheng
Wang, Xintao
Zhou, Jiantao
Qiao, Yu
Dong, Chao
contents Image processing is a fundamental task in computer vision, which aims at enhancing image quality and extracting essential features for subsequent vision applications. Traditionally, task-specific models are developed for individual tasks and designing such models requires distinct expertise. Building upon the success of large language models (LLMs) in natural language processing (NLP), there is a similar trend in computer vision, which focuses on developing large-scale models through pretraining and in-context learning. This paradigm shift reduces the reliance on task-specific models, yielding a powerful unified model to deal with various tasks. However, these advances have predominantly concentrated on high-level vision tasks, with less attention paid to low-level vision tasks. To address this issue, we propose a universal model for general image processing that covers image restoration, image enhancement, image feature extraction tasks, etc. Our proposed framework, named PromptGIP, unifies these diverse image processing tasks within a universal framework. Inspired by NLP question answering (QA) techniques, we employ a visual prompting question answering paradigm. Specifically, we treat the input-output image pair as a structured question-answer sentence, thereby reprogramming the image processing task as a prompting QA problem. PromptGIP can undertake diverse cross-domain tasks using provided visual prompts, eliminating the need for task-specific finetuning. Our methodology offers a universal and adaptive solution to general image processing. While PromptGIP has demonstrated a certain degree of out-of-domain task generalization capability, further research is expected to fully explore its more powerful emergent generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2310_10513
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Unifying Image Processing as Visual Prompting Question Answering
Liu, Yihao
Chen, Xiangyu
Ma, Xianzheng
Wang, Xintao
Zhou, Jiantao
Qiao, Yu
Dong, Chao
Computer Vision and Pattern Recognition
Image and Video Processing
Image processing is a fundamental task in computer vision, which aims at enhancing image quality and extracting essential features for subsequent vision applications. Traditionally, task-specific models are developed for individual tasks and designing such models requires distinct expertise. Building upon the success of large language models (LLMs) in natural language processing (NLP), there is a similar trend in computer vision, which focuses on developing large-scale models through pretraining and in-context learning. This paradigm shift reduces the reliance on task-specific models, yielding a powerful unified model to deal with various tasks. However, these advances have predominantly concentrated on high-level vision tasks, with less attention paid to low-level vision tasks. To address this issue, we propose a universal model for general image processing that covers image restoration, image enhancement, image feature extraction tasks, etc. Our proposed framework, named PromptGIP, unifies these diverse image processing tasks within a universal framework. Inspired by NLP question answering (QA) techniques, we employ a visual prompting question answering paradigm. Specifically, we treat the input-output image pair as a structured question-answer sentence, thereby reprogramming the image processing task as a prompting QA problem. PromptGIP can undertake diverse cross-domain tasks using provided visual prompts, eliminating the need for task-specific finetuning. Our methodology offers a universal and adaptive solution to general image processing. While PromptGIP has demonstrated a certain degree of out-of-domain task generalization capability, further research is expected to fully explore its more powerful emergent generalization.
title Unifying Image Processing as Visual Prompting Question Answering
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2310.10513