Vision Foundry: A System for Training Foundational Vision AI Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Gokmen, Mahmut S., Klusty, Mitchell A., Damron, Evan W., Logan, W. Vaiden, Mullen, Aaron D., Leach, Caroline N., Collier, Emily B., Armstrong, Samuel E., Bumgardner, V. K. Cody
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917143899013120
author Gokmen, Mahmut S.
Klusty, Mitchell A.
Damron, Evan W.
Logan, W. Vaiden
Mullen, Aaron D.
Leach, Caroline N.
Collier, Emily B.
Armstrong, Samuel E.
Bumgardner, V. K. Cody
author_facet Gokmen, Mahmut S.
Klusty, Mitchell A.
Damron, Evan W.
Logan, W. Vaiden
Mullen, Aaron D.
Leach, Caroline N.
Collier, Emily B.
Armstrong, Samuel E.
Bumgardner, V. K. Cody
contents Self-supervised learning (SSL) leverages vast unannotated medical datasets, yet steep technical barriers limit adoption by clinical researchers. We introduce Vision Foundry, a code-free, HIPAA-compliant platform that democratizes pre-training, adaptation, and deployment of foundational vision models. The system integrates the DINO-MX framework, abstracting distributed infrastructure complexities while implementing specialized strategies like Magnification-Aware Distillation (MAD) and Parameter-Efficient Fine-Tuning (PEFT). We validate the platform across domains, including neuropathology segmentation, lung cellularity estimation, and coronary calcium scoring. Our experiments demonstrate that models trained via Vision Foundry significantly outperform generic baselines in segmentation fidelity and regression accuracy, while exhibiting robust zero-shot generalization across imaging protocols. By bridging the gap between advanced representation learning and practical application, Vision Foundry enables domain experts to develop state-of-the-art clinical AI tools with minimal annotation overhead, shifting focus from engineering optimization to clinical discovery.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11837
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vision Foundry: A System for Training Foundational Vision AI Models
Gokmen, Mahmut S.
Klusty, Mitchell A.
Damron, Evan W.
Logan, W. Vaiden
Mullen, Aaron D.
Leach, Caroline N.
Collier, Emily B.
Armstrong, Samuel E.
Bumgardner, V. K. Cody
Quantitative Methods
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Self-supervised learning (SSL) leverages vast unannotated medical datasets, yet steep technical barriers limit adoption by clinical researchers. We introduce Vision Foundry, a code-free, HIPAA-compliant platform that democratizes pre-training, adaptation, and deployment of foundational vision models. The system integrates the DINO-MX framework, abstracting distributed infrastructure complexities while implementing specialized strategies like Magnification-Aware Distillation (MAD) and Parameter-Efficient Fine-Tuning (PEFT). We validate the platform across domains, including neuropathology segmentation, lung cellularity estimation, and coronary calcium scoring. Our experiments demonstrate that models trained via Vision Foundry significantly outperform generic baselines in segmentation fidelity and regression accuracy, while exhibiting robust zero-shot generalization across imaging protocols. By bridging the gap between advanced representation learning and practical application, Vision Foundry enables domain experts to develop state-of-the-art clinical AI tools with minimal annotation overhead, shifting focus from engineering optimization to clinical discovery.
title Vision Foundry: A System for Training Foundational Vision AI Models
topic Quantitative Methods
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2512.11837