From Simple to Professional: A Combinatorial Controllable Image Captioning Agent

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xinran, Diao, Muxi, Li, Baoteng, Zhang, Haiwen, Liang, Kongming, Ma, Zhanyu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913643455578112
author Wang, Xinran
Diao, Muxi
Li, Baoteng
Zhang, Haiwen
Liang, Kongming
Ma, Zhanyu
author_facet Wang, Xinran
Diao, Muxi
Li, Baoteng
Zhang, Haiwen
Liang, Kongming
Ma, Zhanyu
contents The Controllable Image Captioning Agent (CapAgent) is an innovative system designed to bridge the gap between user simplicity and professional-level outputs in image captioning tasks. CapAgent automatically transforms user-provided simple instructions into detailed, professional instructions, enabling precise and context-aware caption generation. By leveraging multimodal large language models (MLLMs) and external tools such as object detection tool and search engines, the system ensures that captions adhere to specified guidelines, including sentiment, keywords, focus, and formatting. CapAgent transparently controls each step of the captioning process, and showcases its reasoning and tool usage at every step, fostering user trust and engagement. The project code is available at https://github.com/xin-ran-w/CapAgent.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11025
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle From Simple to Professional: A Combinatorial Controllable Image Captioning Agent
Wang, Xinran
Diao, Muxi
Li, Baoteng
Zhang, Haiwen
Liang, Kongming
Ma, Zhanyu
Computer Vision and Pattern Recognition
Artificial Intelligence
The Controllable Image Captioning Agent (CapAgent) is an innovative system designed to bridge the gap between user simplicity and professional-level outputs in image captioning tasks. CapAgent automatically transforms user-provided simple instructions into detailed, professional instructions, enabling precise and context-aware caption generation. By leveraging multimodal large language models (MLLMs) and external tools such as object detection tool and search engines, the system ensures that captions adhere to specified guidelines, including sentiment, keywords, focus, and formatting. CapAgent transparently controls each step of the captioning process, and showcases its reasoning and tool usage at every step, fostering user trust and engagement. The project code is available at https://github.com/xin-ran-w/CapAgent.
title From Simple to Professional: A Combinatorial Controllable Image Captioning Agent
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2412.11025