Large-scale entity resolution via microclustering Ewens--Pitman random partitions
Fuente:
arXiv
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915407796895744 |
|---|---|
| author | Beraha, Mario Favaro, Stefano |
| author_facet | Beraha, Mario Favaro, Stefano |
| contents | We introduce the microclustering Ewens--Pitman model for random partitions, obtained by scaling the strength parameter of the Ewens--Pitman model linearly with the sample size. The resulting random partition is shown to have the microclustering property, namely: the size of the largest cluster grows sub-linearly with the sample size, while the number of clusters grows linearly. By leveraging the interplay between the Ewens--Pitman random partition with the Pitman--Yor process, we develop efficient variational inference schemes for posterior computation in entity resolution. Our approach achieves a speed-up of three orders of magnitude over existing Bayesian methods for entity resolution, while maintaining competitive empirical performance. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_18101 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Large-scale entity resolution via microclustering Ewens--Pitman random partitions Beraha, Mario Favaro, Stefano Methodology Statistics Theory Computation Machine Learning We introduce the microclustering Ewens--Pitman model for random partitions, obtained by scaling the strength parameter of the Ewens--Pitman model linearly with the sample size. The resulting random partition is shown to have the microclustering property, namely: the size of the largest cluster grows sub-linearly with the sample size, while the number of clusters grows linearly. By leveraging the interplay between the Ewens--Pitman random partition with the Pitman--Yor process, we develop efficient variational inference schemes for posterior computation in entity resolution. Our approach achieves a speed-up of three orders of magnitude over existing Bayesian methods for entity resolution, while maintaining competitive empirical performance. |
| title | Large-scale entity resolution via microclustering Ewens--Pitman random partitions |
| topic | Methodology Statistics Theory Computation Machine Learning |
| url | https://arxiv.org/abs/2507.18101 |