Autonomous Computer Vision Development with Agentic AI

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Jin, Wahi-Anwa, Muhammad, Park, Sangyun, Shin, Shawn, Hoffman, John M., Brown, Matthew S.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912441351274496
author Kim, Jin
Wahi-Anwa, Muhammad
Park, Sangyun
Shin, Shawn
Hoffman, John M.
Brown, Matthew S.
author_facet Kim, Jin
Wahi-Anwa, Muhammad
Park, Sangyun
Shin, Shawn
Hoffman, John M.
Brown, Matthew S.
contents Agentic Artificial Intelligence (AI) systems leveraging Large Language Models (LLMs) exhibit significant potential for complex reasoning, planning, and tool utilization. We demonstrate that a specialized computer vision system can be built autonomously from a natural language prompt using Agentic AI methods. This involved extending SimpleMind (SM), an open-source Cognitive AI environment with configurable tools for medical image analysis, with an LLM-based agent, implemented using OpenManus, to automate the planning (tool configuration) for a particular computer vision task. We provide a proof-of-concept demonstration that an agentic system can interpret a computer vision task prompt, plan a corresponding SimpleMind workflow by decomposing the task and configuring appropriate tools. From the user input prompt, "provide sm (SimpleMind) config for lungs, heart, and ribs segmentation for cxr (chest x-ray)"), the agent LLM was able to generate the plan (tool configuration file in YAML format), and execute SM-Learn (training) and SM-Think (inference) scripts autonomously. The computer vision agent automatically configured, trained, and tested itself on 50 chest x-ray images, achieving mean dice scores of 0.96, 0.82, 0.83, for lungs, heart, and ribs, respectively. This work shows the potential for autonomous planning and tool configuration that has traditionally been performed by a data scientist in the development of computer vision applications.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11140
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Autonomous Computer Vision Development with Agentic AI
Kim, Jin
Wahi-Anwa, Muhammad
Park, Sangyun
Shin, Shawn
Hoffman, John M.
Brown, Matthew S.
Computer Vision and Pattern Recognition
Artificial Intelligence
Multiagent Systems
Agentic Artificial Intelligence (AI) systems leveraging Large Language Models (LLMs) exhibit significant potential for complex reasoning, planning, and tool utilization. We demonstrate that a specialized computer vision system can be built autonomously from a natural language prompt using Agentic AI methods. This involved extending SimpleMind (SM), an open-source Cognitive AI environment with configurable tools for medical image analysis, with an LLM-based agent, implemented using OpenManus, to automate the planning (tool configuration) for a particular computer vision task. We provide a proof-of-concept demonstration that an agentic system can interpret a computer vision task prompt, plan a corresponding SimpleMind workflow by decomposing the task and configuring appropriate tools. From the user input prompt, "provide sm (SimpleMind) config for lungs, heart, and ribs segmentation for cxr (chest x-ray)"), the agent LLM was able to generate the plan (tool configuration file in YAML format), and execute SM-Learn (training) and SM-Think (inference) scripts autonomously. The computer vision agent automatically configured, trained, and tested itself on 50 chest x-ray images, achieving mean dice scores of 0.96, 0.82, 0.83, for lungs, heart, and ribs, respectively. This work shows the potential for autonomous planning and tool configuration that has traditionally been performed by a data scientist in the development of computer vision applications.
title Autonomous Computer Vision Development with Agentic AI
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multiagent Systems
url https://arxiv.org/abs/2506.11140