Saved in:
Bibliographic Details
Main Authors: Poore, Tyler J, Pinard, Christopher J, Shabbir, Aleena, Lagree, Andrew, Telfer, Andre, Wu, Kuan-Chuen
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.01224
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908572780068864
author Poore, Tyler J
Pinard, Christopher J
Shabbir, Aleena
Lagree, Andrew
Telfer, Andre
Wu, Kuan-Chuen
author_facet Poore, Tyler J
Pinard, Christopher J
Shabbir, Aleena
Lagree, Andrew
Telfer, Andre
Wu, Kuan-Chuen
contents Large language models (LLMs) are increasingly used in clinical settings, yet their performance in veterinary medicine remains underexplored. We evaluated three commercially available veterinary-focused LLM summarization tools (Product 1 [Hachiko] and Products 2 and 3) on a standardized dataset of veterinary oncology records. Using a rubric-guided LLM-as-a-judge framework, summaries were scored across five domains: Factual Accuracy, Completeness, Chronological Order, Clinical Relevance, and Organization. Product 1 achieved the highest overall performance, with a median average score of 4.61 (IQR: 0.73), compared to 2.55 (IQR: 0.78) for Product 2 and 2.45 (IQR: 0.92) for Product 3. It also received perfect median scores in Factual Accuracy and Chronological Order. To assess the internal consistency of the grading framework itself, we repeated the evaluation across three independent runs. The LLM grader demonstrated high reproducibility, with Average Score standard deviations of 0.015 (Product 1), 0.088 (Product 2), and 0.034 (Product 3). These findings highlight the importance of veterinary-specific commercial LLM tools and demonstrate that LLM-as-a-judge evaluation is a scalable and reproducible method for assessing clinical NLP summarization in veterinary medicine.
format Preprint
id arxiv_https___arxiv_org_abs_2510_01224
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Context Matters: Comparison of commercial large language tools in veterinary medicine
Poore, Tyler J
Pinard, Christopher J
Shabbir, Aleena
Lagree, Andrew
Telfer, Andre
Wu, Kuan-Chuen
Computation and Language
Artificial Intelligence
Large language models (LLMs) are increasingly used in clinical settings, yet their performance in veterinary medicine remains underexplored. We evaluated three commercially available veterinary-focused LLM summarization tools (Product 1 [Hachiko] and Products 2 and 3) on a standardized dataset of veterinary oncology records. Using a rubric-guided LLM-as-a-judge framework, summaries were scored across five domains: Factual Accuracy, Completeness, Chronological Order, Clinical Relevance, and Organization. Product 1 achieved the highest overall performance, with a median average score of 4.61 (IQR: 0.73), compared to 2.55 (IQR: 0.78) for Product 2 and 2.45 (IQR: 0.92) for Product 3. It also received perfect median scores in Factual Accuracy and Chronological Order. To assess the internal consistency of the grading framework itself, we repeated the evaluation across three independent runs. The LLM grader demonstrated high reproducibility, with Average Score standard deviations of 0.015 (Product 1), 0.088 (Product 2), and 0.034 (Product 3). These findings highlight the importance of veterinary-specific commercial LLM tools and demonstrate that LLM-as-a-judge evaluation is a scalable and reproducible method for assessing clinical NLP summarization in veterinary medicine.
title Context Matters: Comparison of commercial large language tools in veterinary medicine
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.01224