Region-Normalized DPO for Medical Image Segmentation under Noisy Judges

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kalisch, Hamza, Seibold, Constantin, Kleesiek, Jens, Herrmann, Ken, Jonske, Frederic
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917235572867072
author Kalisch, Hamza
Seibold, Constantin
Kleesiek, Jens
Herrmann, Ken
Jonske, Frederic
author_facet Kalisch, Hamza
Seibold, Constantin
Kleesiek, Jens
Herrmann, Ken
Jonske, Frederic
contents While dense pixel-wise annotations remain the gold standard for medical image segmentation, they are costly to obtain and limit scalability. In contrast, many deployed systems already produce inexpensive automatic quality-control (QC) signals like model agreement, uncertainty measures, or learned mask-quality scores which can be used for further model training without additional ground-truth annotation. However, these signals can be noisy and biased, making preference-based fine-tuning susceptible to harmful updates. We study Direct Preference Optimization (DPO) for segmentation from such noisy judges using proposals generated by a supervised base segmenter trained on a small labeled set. We find that outcomes depend strongly on how preference pairs are mined: selecting the judge's top-ranked proposal can improve peak performance when the judge is reliable, but can amplify harmful errors under weaker judges. We propose Region-Normalized DPO (RN-DPO), a segmentation-aware objective which normalizes preference updates by the size of the disagreement region between masks, reducing the leverage of harmful comparisons and improving optimization stability. Across two medical datasets and multiple regimes, RN-DPO improves sustained performance and stabilizes preference-based fine-tuning, outperforming standard DPO and strong baselines without requiring additional pixel annotations.
format Preprint
id arxiv_https___arxiv_org_abs_2601_23222
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Region-Normalized DPO for Medical Image Segmentation under Noisy Judges
Kalisch, Hamza
Seibold, Constantin
Kleesiek, Jens
Herrmann, Ken
Jonske, Frederic
Computer Vision and Pattern Recognition
While dense pixel-wise annotations remain the gold standard for medical image segmentation, they are costly to obtain and limit scalability. In contrast, many deployed systems already produce inexpensive automatic quality-control (QC) signals like model agreement, uncertainty measures, or learned mask-quality scores which can be used for further model training without additional ground-truth annotation. However, these signals can be noisy and biased, making preference-based fine-tuning susceptible to harmful updates. We study Direct Preference Optimization (DPO) for segmentation from such noisy judges using proposals generated by a supervised base segmenter trained on a small labeled set. We find that outcomes depend strongly on how preference pairs are mined: selecting the judge's top-ranked proposal can improve peak performance when the judge is reliable, but can amplify harmful errors under weaker judges. We propose Region-Normalized DPO (RN-DPO), a segmentation-aware objective which normalizes preference updates by the size of the disagreement region between masks, reducing the leverage of harmful comparisons and improving optimization stability. Across two medical datasets and multiple regimes, RN-DPO improves sustained performance and stabilizes preference-based fine-tuning, outperforming standard DPO and strong baselines without requiring additional pixel annotations.
title Region-Normalized DPO for Medical Image Segmentation under Noisy Judges
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.23222