Filtering out mislabeled training instances using black-box optimization and quantum annealing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Otsuka, Makoto, Kodama, Kento, Morita, Keisuke, Ohzeki, Masayuki
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911202020425728
author Otsuka, Makoto
Kodama, Kento
Morita, Keisuke
Ohzeki, Masayuki
author_facet Otsuka, Makoto
Kodama, Kento
Morita, Keisuke
Ohzeki, Masayuki
contents This study proposes an approach for removing mislabeled instances from contaminated training datasets by combining surrogate model-based black-box optimization (BBO) with postprocessing and quantum annealing. Mislabeled training instances, a common issue in real-world datasets, often degrade model generalization, necessitating robust and efficient noise-removal strategies. The proposed method evaluates filtered training subsets based on validation loss, iteratively refines loss estimates through surrogate model-based BBO with postprocessing, and leverages quantum annealing to efficiently sample diverse training subsets with low validation error. Experiments on a noisy majority bit task demonstrate the method's ability to prioritize the removal of high-risk mislabeled instances. Integrating D-Wave's clique sampler running on a physical quantum annealer achieves faster optimization and higher-quality training subsets compared to OpenJij's simulated quantum annealing sampler or Neal's simulated annealing sampler, offering a scalable framework for enhancing dataset quality. This work highlights the effectiveness of the proposed method for supervised learning tasks, with future directions including its application to unsupervised learning, real-world datasets, and large-scale implementations.
format Preprint
id arxiv_https___arxiv_org_abs_2501_06916
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Filtering out mislabeled training instances using black-box optimization and quantum annealing
Otsuka, Makoto
Kodama, Kento
Morita, Keisuke
Ohzeki, Masayuki
Machine Learning
Statistical Mechanics
Quantum Physics
This study proposes an approach for removing mislabeled instances from contaminated training datasets by combining surrogate model-based black-box optimization (BBO) with postprocessing and quantum annealing. Mislabeled training instances, a common issue in real-world datasets, often degrade model generalization, necessitating robust and efficient noise-removal strategies. The proposed method evaluates filtered training subsets based on validation loss, iteratively refines loss estimates through surrogate model-based BBO with postprocessing, and leverages quantum annealing to efficiently sample diverse training subsets with low validation error. Experiments on a noisy majority bit task demonstrate the method's ability to prioritize the removal of high-risk mislabeled instances. Integrating D-Wave's clique sampler running on a physical quantum annealer achieves faster optimization and higher-quality training subsets compared to OpenJij's simulated quantum annealing sampler or Neal's simulated annealing sampler, offering a scalable framework for enhancing dataset quality. This work highlights the effectiveness of the proposed method for supervised learning tasks, with future directions including its application to unsupervised learning, real-world datasets, and large-scale implementations.
title Filtering out mislabeled training instances using black-box optimization and quantum annealing
topic Machine Learning
Statistical Mechanics
Quantum Physics
url https://arxiv.org/abs/2501.06916