TOPJoin: A Context-Aware Multi-Criteria Approach for Joinable Column Search
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913942821928960 |
|---|---|
| author | Kokel, Harsha Khatiwada, Aamod Pedapati, Tejaswini Ananthakrishnan, Haritha Hassanzadeh, Oktie Samulowitz, Horst Srinivas, Kavitha |
| author_facet | Kokel, Harsha Khatiwada, Aamod Pedapati, Tejaswini Ananthakrishnan, Haritha Hassanzadeh, Oktie Samulowitz, Horst Srinivas, Kavitha |
| contents | One of the major challenges in enterprise data analysis is the task of finding joinable tables that are conceptually related and provide meaningful insights. Traditionally, joinable tables have been discovered through a search for similar columns, where two columns are considered similar syntactically if there is a set overlap or they are considered similar semantically if either the column embeddings or value embeddings are closer in the embedding space. However, for enterprise data lakes, column similarity is not sufficient to identify joinable columns and tables. The context of the query column is important. Hence, in this work, we first define context-aware column joinability. Then we propose a multi-criteria approach, called TOPJoin, for joinable column search. We evaluate TOPJoin against existing join search baselines over one academic and one real-world join search benchmark. Through experiments, we find that TOPJoin performs better on both benchmarks than the baselines. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_11505 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | TOPJoin: A Context-Aware Multi-Criteria Approach for Joinable Column Search Kokel, Harsha Khatiwada, Aamod Pedapati, Tejaswini Ananthakrishnan, Haritha Hassanzadeh, Oktie Samulowitz, Horst Srinivas, Kavitha Databases One of the major challenges in enterprise data analysis is the task of finding joinable tables that are conceptually related and provide meaningful insights. Traditionally, joinable tables have been discovered through a search for similar columns, where two columns are considered similar syntactically if there is a set overlap or they are considered similar semantically if either the column embeddings or value embeddings are closer in the embedding space. However, for enterprise data lakes, column similarity is not sufficient to identify joinable columns and tables. The context of the query column is important. Hence, in this work, we first define context-aware column joinability. Then we propose a multi-criteria approach, called TOPJoin, for joinable column search. We evaluate TOPJoin against existing join search baselines over one academic and one real-world join search benchmark. Through experiments, we find that TOPJoin performs better on both benchmarks than the baselines. |
| title | TOPJoin: A Context-Aware Multi-Criteria Approach for Joinable Column Search |
| topic | Databases |
| url | https://arxiv.org/abs/2507.11505 |