Evaluating Joinable Column Discovery Approaches for Context-Aware Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kokel, Harsha, Khatiwada, Aamod, Pedapati, Tejaswini, Ananthakrishnan, Haritha, Hassanzadeh, Oktie, Samulowitz, Horst, Srinivas, Kavitha
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917048199675904
author Kokel, Harsha
Khatiwada, Aamod
Pedapati, Tejaswini
Ananthakrishnan, Haritha
Hassanzadeh, Oktie
Samulowitz, Horst
Srinivas, Kavitha
author_facet Kokel, Harsha
Khatiwada, Aamod
Pedapati, Tejaswini
Ananthakrishnan, Haritha
Hassanzadeh, Oktie
Samulowitz, Horst
Srinivas, Kavitha
contents Joinable Column Discovery is a critical challenge in automating enterprise data analysis. While existing approaches focus on syntactic overlap and semantic similarity, there remains limited understanding of which methods perform best for different data characteristics and how multiple criteria influence discovery effectiveness. We present a comprehensive experimental evaluation of joinable column discovery methods across diverse scenarios. Our study compares syntactic and semantic techniques on seven benchmarks covering relational databases and data lakes. We analyze six key criteria -- unique values, intersection size, join size, reverse join size, value semantics, and metadata semantics -- and examine how combining them through ensemble ranking affects performance. Our analysis reveals differences in method behavior across data contexts and highlights the benefits of integrating multiple criteria for robust join discovery. We provide empirical evidence on when each criterion matters, compare pre-trained embedding models for semantic joins, and offer practical guidelines for selecting suitable methods based on dataset characteristics. Our findings show that metadata and value semantics are crucial for data lakes, size-based criteria play a stronger role in relational databases, and ensemble approaches consistently outperform single-criterion methods.
format Preprint
id arxiv_https___arxiv_org_abs_2510_24599
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Joinable Column Discovery Approaches for Context-Aware Search
Kokel, Harsha
Khatiwada, Aamod
Pedapati, Tejaswini
Ananthakrishnan, Haritha
Hassanzadeh, Oktie
Samulowitz, Horst
Srinivas, Kavitha
Databases
Joinable Column Discovery is a critical challenge in automating enterprise data analysis. While existing approaches focus on syntactic overlap and semantic similarity, there remains limited understanding of which methods perform best for different data characteristics and how multiple criteria influence discovery effectiveness. We present a comprehensive experimental evaluation of joinable column discovery methods across diverse scenarios. Our study compares syntactic and semantic techniques on seven benchmarks covering relational databases and data lakes. We analyze six key criteria -- unique values, intersection size, join size, reverse join size, value semantics, and metadata semantics -- and examine how combining them through ensemble ranking affects performance. Our analysis reveals differences in method behavior across data contexts and highlights the benefits of integrating multiple criteria for robust join discovery. We provide empirical evidence on when each criterion matters, compare pre-trained embedding models for semantic joins, and offer practical guidelines for selecting suitable methods based on dataset characteristics. Our findings show that metadata and value semantics are crucial for data lakes, size-based criteria play a stronger role in relational databases, and ensemble approaches consistently outperform single-criterion methods.
title Evaluating Joinable Column Discovery Approaches for Context-Aware Search
topic Databases
url https://arxiv.org/abs/2510.24599