ChatGPT and general-purpose AI count fruits in pictures surprisingly well

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mengsuwan, Konlavach, Palacio, Juan Camilo Rivera, Ryo, Masahiro
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916203266572288
author Mengsuwan, Konlavach
Palacio, Juan Camilo Rivera
Ryo, Masahiro
author_facet Mengsuwan, Konlavach
Palacio, Juan Camilo Rivera
Ryo, Masahiro
contents Object counting is a popular task in deep learning applications in various domains, including agriculture. A conventional deep learning approach requires a large amount of training data, often a logistic problem in a real-world application. To address this issue, we examined how well ChatGPT (GPT4V) and a general-purpose AI (foundation model for object counting, T-Rex) can count the number of fruit bodies (coffee cherries) in 100 images. The foundation model with few-shot learning outperformed the trained YOLOv8 model (R2 = 0.923 and 0.900, respectively). ChatGPT also showed some interesting potential, especially when few-shot learning with human feedback was applied (R2 = 0.360 and 0.460, respectively). Moreover, we examined the time required for implementation as a practical question. Obtaining the results with the foundation model and ChatGPT were much shorter than the YOLOv8 model (0.83 hrs, 1.75 hrs, and 161 hrs). We interpret these results as two surprises for deep learning users in applied domains: a foundation model with few-shot domain-specific learning can drastically save time and effort compared to the conventional approach, and ChatGPT can reveal a relatively good performance. Both approaches do not need coding skills, which can foster AI education and dissemination.
format Preprint
id arxiv_https___arxiv_org_abs_2404_08515
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ChatGPT and general-purpose AI count fruits in pictures surprisingly well
Mengsuwan, Konlavach
Palacio, Juan Camilo Rivera
Ryo, Masahiro
Computer Vision and Pattern Recognition
Image and Video Processing
Object counting is a popular task in deep learning applications in various domains, including agriculture. A conventional deep learning approach requires a large amount of training data, often a logistic problem in a real-world application. To address this issue, we examined how well ChatGPT (GPT4V) and a general-purpose AI (foundation model for object counting, T-Rex) can count the number of fruit bodies (coffee cherries) in 100 images. The foundation model with few-shot learning outperformed the trained YOLOv8 model (R2 = 0.923 and 0.900, respectively). ChatGPT also showed some interesting potential, especially when few-shot learning with human feedback was applied (R2 = 0.360 and 0.460, respectively). Moreover, we examined the time required for implementation as a practical question. Obtaining the results with the foundation model and ChatGPT were much shorter than the YOLOv8 model (0.83 hrs, 1.75 hrs, and 161 hrs). We interpret these results as two surprises for deep learning users in applied domains: a foundation model with few-shot domain-specific learning can drastically save time and effort compared to the conventional approach, and ChatGPT can reveal a relatively good performance. Both approaches do not need coding skills, which can foster AI education and dissemination.
title ChatGPT and general-purpose AI count fruits in pictures surprisingly well
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2404.08515