Building Trust in Virtual Immunohistochemistry: Automated Assessment of Image Quality

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kataria, Tushar, Dubey, Shikha, Bronner, Mary, Jedrzkiewicz, Jolanta, Brintz, Ben J., Elhabian, Shireen Y., Knudsen, Beatrice S.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911251974586368
author Kataria, Tushar
Dubey, Shikha
Bronner, Mary
Jedrzkiewicz, Jolanta
Brintz, Ben J.
Elhabian, Shireen Y.
Knudsen, Beatrice S.
author_facet Kataria, Tushar
Dubey, Shikha
Bronner, Mary
Jedrzkiewicz, Jolanta
Brintz, Ben J.
Elhabian, Shireen Y.
Knudsen, Beatrice S.
contents Deep learning models can generate virtual immunohistochemistry (IHC) stains from hematoxylin and eosin (H&E) images, offering a scalable and low-cost alternative to laboratory IHC. However, reliable evaluation of image quality remains a challenge as current texture- and distribution-based metrics quantify image fidelity rather than the accuracy of IHC staining. Here, we introduce an automated and accuracy grounded framework to determine image quality across sixteen paired or unpaired image translation models. Using color deconvolution, we generate masks of pixels stained brown (i.e., IHC-positive) as predicted by each virtual IHC model. We use the segmented masks of real and virtual IHC to compute stain accuracy metrics (Dice, IoU, Hausdorff distance) that directly quantify correct pixel - level labeling without needing expert manual annotations. Our results demonstrate that conventional image fidelity metrics, including Frechet Inception Distance (FID), peak signal-to-noise ratio (PSNR), and structural similarity (SSIM), correlate poorly with stain accuracy and pathologist assessment. Paired models such as PyramidPix2Pix and AdaptiveNCE achieve the highest stain accuracy, whereas unpaired diffusion- and GAN-based models are less reliable in providing accurate IHC positive pixel labels. Moreover, whole-slide images (WSI) reveal performance declines that are invisible in patch-based evaluations, emphasizing the need for WSI-level benchmarks. Together, this framework defines a reproducible approach for assessing the quality of virtual IHC models, a critical step to accelerate translation towards routine use by pathologists.
format Preprint
id arxiv_https___arxiv_org_abs_2511_04615
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Building Trust in Virtual Immunohistochemistry: Automated Assessment of Image Quality
Kataria, Tushar
Dubey, Shikha
Bronner, Mary
Jedrzkiewicz, Jolanta
Brintz, Ben J.
Elhabian, Shireen Y.
Knudsen, Beatrice S.
Computer Vision and Pattern Recognition
Deep learning models can generate virtual immunohistochemistry (IHC) stains from hematoxylin and eosin (H&E) images, offering a scalable and low-cost alternative to laboratory IHC. However, reliable evaluation of image quality remains a challenge as current texture- and distribution-based metrics quantify image fidelity rather than the accuracy of IHC staining. Here, we introduce an automated and accuracy grounded framework to determine image quality across sixteen paired or unpaired image translation models. Using color deconvolution, we generate masks of pixels stained brown (i.e., IHC-positive) as predicted by each virtual IHC model. We use the segmented masks of real and virtual IHC to compute stain accuracy metrics (Dice, IoU, Hausdorff distance) that directly quantify correct pixel - level labeling without needing expert manual annotations. Our results demonstrate that conventional image fidelity metrics, including Frechet Inception Distance (FID), peak signal-to-noise ratio (PSNR), and structural similarity (SSIM), correlate poorly with stain accuracy and pathologist assessment. Paired models such as PyramidPix2Pix and AdaptiveNCE achieve the highest stain accuracy, whereas unpaired diffusion- and GAN-based models are less reliable in providing accurate IHC positive pixel labels. Moreover, whole-slide images (WSI) reveal performance declines that are invisible in patch-based evaluations, emphasizing the need for WSI-level benchmarks. Together, this framework defines a reproducible approach for assessing the quality of virtual IHC models, a critical step to accelerate translation towards routine use by pathologists.
title Building Trust in Virtual Immunohistochemistry: Automated Assessment of Image Quality
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.04615