Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Thanh Tam, Ren, Zhao, Pham, Trinh, Huynh, Thanh Trung, Nguyen, Phi Le, Yin, Hongzhi, Nguyen, Quoc Viet Hung
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912128419495936
author Nguyen, Thanh Tam
Ren, Zhao
Pham, Trinh
Huynh, Thanh Trung
Nguyen, Phi Le
Yin, Hongzhi
Nguyen, Quoc Viet Hung
author_facet Nguyen, Thanh Tam
Ren, Zhao
Pham, Trinh
Huynh, Thanh Trung
Nguyen, Phi Le
Yin, Hongzhi
Nguyen, Quoc Viet Hung
contents The rapid advancement of large language models (LLMs) and multimodal learning has transformed digital content creation and manipulation. Traditional visual editing tools require significant expertise, limiting accessibility. Recent strides in instruction-based editing have enabled intuitive interaction with visual content, using natural language as a bridge between user intent and complex editing operations. This survey provides an overview of these techniques, focusing on how LLMs and multimodal models empower users to achieve precise visual modifications without deep technical knowledge. By synthesizing over 100 publications, we explore methods from generative adversarial networks to diffusion models, examining multimodal integration for fine-grained content control. We discuss practical applications across domains such as fashion, 3D scene manipulation, and video synthesis, highlighting increased accessibility and alignment with human intuition. Our survey compares existing literature, emphasizing LLM-empowered editing, and identifies key challenges to stimulate further research. We aim to democratize powerful visual editing across various industries, from entertainment to education. Interested readers are encouraged to access our repository at https://github.com/tamlhp/awesome-instruction-editing.
format Preprint
id arxiv_https___arxiv_org_abs_2411_09955
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era
Nguyen, Thanh Tam
Ren, Zhao
Pham, Trinh
Huynh, Thanh Trung
Nguyen, Phi Le
Yin, Hongzhi
Nguyen, Quoc Viet Hung
Computer Vision and Pattern Recognition
Artificial Intelligence
Human-Computer Interaction
Machine Learning
Multimedia
The rapid advancement of large language models (LLMs) and multimodal learning has transformed digital content creation and manipulation. Traditional visual editing tools require significant expertise, limiting accessibility. Recent strides in instruction-based editing have enabled intuitive interaction with visual content, using natural language as a bridge between user intent and complex editing operations. This survey provides an overview of these techniques, focusing on how LLMs and multimodal models empower users to achieve precise visual modifications without deep technical knowledge. By synthesizing over 100 publications, we explore methods from generative adversarial networks to diffusion models, examining multimodal integration for fine-grained content control. We discuss practical applications across domains such as fashion, 3D scene manipulation, and video synthesis, highlighting increased accessibility and alignment with human intuition. Our survey compares existing literature, emphasizing LLM-empowered editing, and identifies key challenges to stimulate further research. We aim to democratize powerful visual editing across various industries, from entertainment to education. Interested readers are encouraged to access our repository at https://github.com/tamlhp/awesome-instruction-editing.
title Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Human-Computer Interaction
Machine Learning
Multimedia
url https://arxiv.org/abs/2411.09955