A shortest-path based clustering algorithm for joint human-machine analysis of complex datasets

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pizzagalli, Diego Ulisse, Gonzalez, Santiago Fernandez, Krause, Rolf
Format: Preprint
Published: 2018
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929517641072640
author Pizzagalli, Diego Ulisse
Gonzalez, Santiago Fernandez
Krause, Rolf
author_facet Pizzagalli, Diego Ulisse
Gonzalez, Santiago Fernandez
Krause, Rolf
contents Clustering is a technique for the analysis of datasets obtained by empirical studies in several disciplines with a major application for biomedical research. Essentially, clustering algorithms are executed by machines aiming at finding groups of related points in a dataset. However, the result of grouping depends on both metrics for point-to-point similarity and rules for point-to-group association. Indeed, non-appropriate metrics and rules can lead to undesirable clustering artifacts. This is especially relevant for datasets, where groups with heterogeneous structures co-exist. In this work, we propose an algorithm that achieves clustering by exploring the paths between points. This allows both, to evaluate the properties of the path (such as gaps, density variations, etc.), and expressing the preference for certain paths. Moreover, our algorithm supports the integration of existing knowledge about admissible and non-admissible clusters by training a path classifier. We demonstrate the accuracy of the proposed method on challenging datasets including points from synthetic shapes in publicly available benchmarks and microscopy data.
format Preprint
id arxiv_https___arxiv_org_abs_1812_11850
institution arXiv
publishDate 2018
record_format arxiv
spellingShingle A shortest-path based clustering algorithm for joint human-machine analysis of complex datasets
Pizzagalli, Diego Ulisse
Gonzalez, Santiago Fernandez
Krause, Rolf
Quantitative Methods
Artificial Intelligence
Human-Computer Interaction
Machine Learning
Clustering is a technique for the analysis of datasets obtained by empirical studies in several disciplines with a major application for biomedical research. Essentially, clustering algorithms are executed by machines aiming at finding groups of related points in a dataset. However, the result of grouping depends on both metrics for point-to-point similarity and rules for point-to-group association. Indeed, non-appropriate metrics and rules can lead to undesirable clustering artifacts. This is especially relevant for datasets, where groups with heterogeneous structures co-exist. In this work, we propose an algorithm that achieves clustering by exploring the paths between points. This allows both, to evaluate the properties of the path (such as gaps, density variations, etc.), and expressing the preference for certain paths. Moreover, our algorithm supports the integration of existing knowledge about admissible and non-admissible clusters by training a path classifier. We demonstrate the accuracy of the proposed method on challenging datasets including points from synthetic shapes in publicly available benchmarks and microscopy data.
title A shortest-path based clustering algorithm for joint human-machine analysis of complex datasets
topic Quantitative Methods
Artificial Intelligence
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/1812.11850