From Simple to Professional: A Combinatorial Controllable Image Captioning Agent
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913643455578112 |
|---|---|
| author | Wang, Xinran Diao, Muxi Li, Baoteng Zhang, Haiwen Liang, Kongming Ma, Zhanyu |
| author_facet | Wang, Xinran Diao, Muxi Li, Baoteng Zhang, Haiwen Liang, Kongming Ma, Zhanyu |
| contents | The Controllable Image Captioning Agent (CapAgent) is an innovative system designed to bridge the gap between user simplicity and professional-level outputs in image captioning tasks. CapAgent automatically transforms user-provided simple instructions into detailed, professional instructions, enabling precise and context-aware caption generation. By leveraging multimodal large language models (MLLMs) and external tools such as object detection tool and search engines, the system ensures that captions adhere to specified guidelines, including sentiment, keywords, focus, and formatting. CapAgent transparently controls each step of the captioning process, and showcases its reasoning and tool usage at every step, fostering user trust and engagement. The project code is available at https://github.com/xin-ran-w/CapAgent. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_11025 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | From Simple to Professional: A Combinatorial Controllable Image Captioning Agent Wang, Xinran Diao, Muxi Li, Baoteng Zhang, Haiwen Liang, Kongming Ma, Zhanyu Computer Vision and Pattern Recognition Artificial Intelligence The Controllable Image Captioning Agent (CapAgent) is an innovative system designed to bridge the gap between user simplicity and professional-level outputs in image captioning tasks. CapAgent automatically transforms user-provided simple instructions into detailed, professional instructions, enabling precise and context-aware caption generation. By leveraging multimodal large language models (MLLMs) and external tools such as object detection tool and search engines, the system ensures that captions adhere to specified guidelines, including sentiment, keywords, focus, and formatting. CapAgent transparently controls each step of the captioning process, and showcases its reasoning and tool usage at every step, fostering user trust and engagement. The project code is available at https://github.com/xin-ran-w/CapAgent. |
| title | From Simple to Professional: A Combinatorial Controllable Image Captioning Agent |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2412.11025 |