Leveraging Large Language Models to Effectively Generate Visual Data for Canine Musculoskeletal Diagnoses

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Thißen, Martin, Tran, Thi Ngoc Diep, Ratsch, Barbara Esteve, Schönbein, Ben Joel, Trapp, Ute, Egner, Beate, Piat, Romana, Hergenröther, Elke
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912589651378176
author Thißen, Martin
Tran, Thi Ngoc Diep
Ratsch, Barbara Esteve
Schönbein, Ben Joel
Trapp, Ute
Egner, Beate
Piat, Romana
Hergenröther, Elke
author_facet Thißen, Martin
Tran, Thi Ngoc Diep
Ratsch, Barbara Esteve
Schönbein, Ben Joel
Trapp, Ute
Egner, Beate
Piat, Romana
Hergenröther, Elke
contents It is well-established that more data generally improves AI model performance. However, data collection can be challenging for certain tasks due to the rarity of occurrences or high costs. These challenges are evident in our use case, where we apply AI models to a novel approach for visually documenting the musculoskeletal condition of dogs. Here, abnormalities are marked as colored strokes on a body map of a dog. Since these strokes correspond to distinct muscles or joints, they can be mapped to the textual domain in which large language models (LLMs) operate. LLMs have demonstrated impressive capabilities across a wide range of tasks, including medical applications, offering promising potential for generating synthetic training data. In this work, we investigate whether LLMs can effectively generate synthetic visual training data for canine musculoskeletal diagnoses. For this, we developed a mapping that segments visual documentations into over 200 labeled regions representing muscles or joints. Using techniques like guided decoding, chain-of-thought reasoning, and few-shot prompting, we generated 1,000 synthetic visual documentations for patellar luxation (kneecap dislocation) diagnosis, the diagnosis for which we have the most real-world data. Our analysis shows that the generated documentations are sensitive to location and severity of the diagnosis while remaining independent of the dog's sex. We further generated 1,000 visual documentations for various other diagnoses to create a binary classification dataset. A model trained solely on this synthetic data achieved an F1 score of 88% on 70 real-world documentations. These results demonstrate the potential of LLM-generated synthetic data, which is particularly valuable for addressing data scarcity in rare diseases. While our methodology is tailored to the medical domain, the insights and techniques can be adapted to other fields.
format Preprint
id arxiv_https___arxiv_org_abs_2509_12866
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Large Language Models to Effectively Generate Visual Data for Canine Musculoskeletal Diagnoses
Thißen, Martin
Tran, Thi Ngoc Diep
Ratsch, Barbara Esteve
Schönbein, Ben Joel
Trapp, Ute
Egner, Beate
Piat, Romana
Hergenröther, Elke
Computer Vision and Pattern Recognition
It is well-established that more data generally improves AI model performance. However, data collection can be challenging for certain tasks due to the rarity of occurrences or high costs. These challenges are evident in our use case, where we apply AI models to a novel approach for visually documenting the musculoskeletal condition of dogs. Here, abnormalities are marked as colored strokes on a body map of a dog. Since these strokes correspond to distinct muscles or joints, they can be mapped to the textual domain in which large language models (LLMs) operate. LLMs have demonstrated impressive capabilities across a wide range of tasks, including medical applications, offering promising potential for generating synthetic training data. In this work, we investigate whether LLMs can effectively generate synthetic visual training data for canine musculoskeletal diagnoses. For this, we developed a mapping that segments visual documentations into over 200 labeled regions representing muscles or joints. Using techniques like guided decoding, chain-of-thought reasoning, and few-shot prompting, we generated 1,000 synthetic visual documentations for patellar luxation (kneecap dislocation) diagnosis, the diagnosis for which we have the most real-world data. Our analysis shows that the generated documentations are sensitive to location and severity of the diagnosis while remaining independent of the dog's sex. We further generated 1,000 visual documentations for various other diagnoses to create a binary classification dataset. A model trained solely on this synthetic data achieved an F1 score of 88% on 70 real-world documentations. These results demonstrate the potential of LLM-generated synthetic data, which is particularly valuable for addressing data scarcity in rare diseases. While our methodology is tailored to the medical domain, the insights and techniques can be adapted to other fields.
title Leveraging Large Language Models to Effectively Generate Visual Data for Canine Musculoskeletal Diagnoses
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.12866