Automatic String Data Validation with Pattern Discovery
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Xinwei, Zhao, Jing, Di, Peng, Xiao, Chuan, Mao, Rui, Ji, Yan, Onizuka, Makoto, Ding, Zishuo, Shang, Weiyi, Qin, Jianbin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Privacy-Enhanced Database Synthesis for Benchmark Publishing (Technical Report)
von: Ge, Yunqing, et al.
Veröffentlicht: (2024)
von: Ge, Yunqing, et al.
Veröffentlicht: (2024)
Ultraverse: A System-Centric Framework for Efficient What-If Analysis for Database-Intensive Web Applications
von: Ko, Ronny, et al.
Veröffentlicht: (2022)
von: Ko, Ronny, et al.
Veröffentlicht: (2022)
ShapleyPipe: Hierarchical Shapley Search for Data Preparation Pipeline Construction
von: Chang, Jing, et al.
Veröffentlicht: (2025)
von: Chang, Jing, et al.
Veröffentlicht: (2025)
LLaPipe: LLM-Guided Reinforcement Learning for Automated Data Preparation Pipeline Construction
von: Chang, Jing, et al.
Veröffentlicht: (2025)
von: Chang, Jing, et al.
Veröffentlicht: (2025)
SoftPipe: A Soft-Guided Reinforcement Learning Framework for Automated Data Preparation
von: Chang, Jing, et al.
Veröffentlicht: (2025)
von: Chang, Jing, et al.
Veröffentlicht: (2025)
Efficient Query Rewrite Rule Discovery via Standardized Enumeration and Learning-to-Rank(extend)
von: Zhang, Yuan, et al.
Veröffentlicht: (2026)
von: Zhang, Yuan, et al.
Veröffentlicht: (2026)
LAKEGEN: A LLM-based Tabular Corpus Generator for Evaluating Dataset Discovery in Data Lakes
von: Dai, Zhenwei, et al.
Veröffentlicht: (2025)
von: Dai, Zhenwei, et al.
Veröffentlicht: (2025)
FeatNavigator: Automatic Feature Augmentation on Tabular Data
von: Liang, Jiaming, et al.
Veröffentlicht: (2024)
von: Liang, Jiaming, et al.
Veröffentlicht: (2024)
OmniMatch: Effective Self-Supervised Any-Join Discovery in Tabular Data Repositories
von: Koutras, Christos, et al.
Veröffentlicht: (2024)
von: Koutras, Christos, et al.
Veröffentlicht: (2024)
Approximate Nearest Neighbor Search for Modern AI: A Projection-Augmented Graph Approach
von: Lu, Kejing, et al.
Veröffentlicht: (2026)
von: Lu, Kejing, et al.
Veröffentlicht: (2026)
Exploring the Heterogeneity of Tabular Data: A Diversity-aware Data Generator via LLMs
von: Tang, Yafeng, et al.
Veröffentlicht: (2025)
von: Tang, Yafeng, et al.
Veröffentlicht: (2025)
The Past Still Matters: A Temporally-Valid Data Discovery System
von: Esmailoghli, Mahdi, et al.
Veröffentlicht: (2025)
von: Esmailoghli, Mahdi, et al.
Veröffentlicht: (2025)
Multiple Index Merge for Approximate Nearest Neighbor Search
von: Jing, Liuchang, et al.
Veröffentlicht: (2026)
von: Jing, Liuchang, et al.
Veröffentlicht: (2026)
Data Overvaluation Attack and Truthful Data Valuation in Federated Learning
von: Zheng, Shuyuan, et al.
Veröffentlicht: (2025)
von: Zheng, Shuyuan, et al.
Veröffentlicht: (2025)
Validating Temporal Compliance Patterns: A Unified Approach with $MTL_f$ over various Data Models
von: Zaki, Nesma M., et al.
Veröffentlicht: (2024)
von: Zaki, Nesma M., et al.
Veröffentlicht: (2024)
Automatic Data Repair: Are We Ready to Deploy?
von: Ni, Wei, et al.
Veröffentlicht: (2023)
von: Ni, Wei, et al.
Veröffentlicht: (2023)
Towards Operationalizing Heterogeneous Data Discovery
von: Wang, Jin, et al.
Veröffentlicht: (2025)
von: Wang, Jin, et al.
Veröffentlicht: (2025)
BOD: Blindly Optimal Data Discovery
von: Hoang, Thomas
Veröffentlicht: (2024)
von: Hoang, Thomas
Veröffentlicht: (2024)
Large Language Models as Data Preprocessors
von: Zhang, Haochen, et al.
Veröffentlicht: (2023)
von: Zhang, Haochen, et al.
Veröffentlicht: (2023)
Efficient Discovery of Significant Patterns with Few-Shot Resampling
von: Pellegrina, Leonardo, et al.
Veröffentlicht: (2024)
von: Pellegrina, Leonardo, et al.
Veröffentlicht: (2024)
Instance-Optimized String Fingerprints
von: Stoian, Mihail, et al.
Veröffentlicht: (2025)
von: Stoian, Mihail, et al.
Veröffentlicht: (2025)
ODIN: A NL2SQL Recommender to Handle Schema Ambiguity
von: Vaidya, Kapil, et al.
Veröffentlicht: (2025)
von: Vaidya, Kapil, et al.
Veröffentlicht: (2025)
Dataset Discovery via Line Charts
von: Ji, Daomin, et al.
Veröffentlicht: (2024)
von: Ji, Daomin, et al.
Veröffentlicht: (2024)
NSPG-Miner: Mining Repetitive Negative Sequential Patterns
von: Li, Yan, et al.
Veröffentlicht: (2025)
von: Li, Yan, et al.
Veröffentlicht: (2025)
Blend: A Unified Data Discovery System
von: Esmailoghli, Mahdi, et al.
Veröffentlicht: (2023)
von: Esmailoghli, Mahdi, et al.
Veröffentlicht: (2023)
FREYJA: Efficient Join Discovery in Data Lakes
von: Maynou, Marc, et al.
Veröffentlicht: (2024)
von: Maynou, Marc, et al.
Veröffentlicht: (2024)
Limitations of Validity Intervals in Data Freshness Management
von: Kang, Kyoung-Don
Veröffentlicht: (2024)
von: Kang, Kyoung-Don
Veröffentlicht: (2024)
SilentWood: Private Inference Over Gradient-Boosting Decision Forests
von: Ko, Ronny, et al.
Veröffentlicht: (2024)
von: Ko, Ronny, et al.
Veröffentlicht: (2024)
Downsizing Diffusion Models for Cardinality Estimation
von: Mu, Xinhe, et al.
Veröffentlicht: (2025)
von: Mu, Xinhe, et al.
Veröffentlicht: (2025)
Birdie: Natural Language-Driven Table Discovery Using Differentiable Search Index
von: Guo, Yuxiang, et al.
Veröffentlicht: (2025)
von: Guo, Yuxiang, et al.
Veröffentlicht: (2025)
Unified Data Discovery across Query Modalities and User Intents
von: Wang, Tingting, et al.
Veröffentlicht: (2026)
von: Wang, Tingting, et al.
Veröffentlicht: (2026)
DataLens: Enhancing Dataset Discovery via Network Topologies
von: Ollagnier, Anaïs, et al.
Veröffentlicht: (2025)
von: Ollagnier, Anaïs, et al.
Veröffentlicht: (2025)
LITS: An Optimized Learned Index for Strings (An Extended Version)
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
On Reporting Durable Patterns in Temporal Proximity Graphs
von: Agarwal, Pankaj K., et al.
Veröffentlicht: (2024)
von: Agarwal, Pankaj K., et al.
Veröffentlicht: (2024)
Automated Data Quality Validation in an End-to-End GNN Framework
von: Dong, Sijie, et al.
Veröffentlicht: (2025)
von: Dong, Sijie, et al.
Veröffentlicht: (2025)
Integrating Knowledge Graphs and Visualization Dashboards for Advance Data Discovery in VESA
von: Betz, Pawandeep Kaur, et al.
Veröffentlicht: (2024)
von: Betz, Pawandeep Kaur, et al.
Veröffentlicht: (2024)
GenIE - Simulator-Driven Iterative Data Exploration for Scientific Discovery
von: Colaco, Ashwin Gerard, et al.
Veröffentlicht: (2025)
von: Colaco, Ashwin Gerard, et al.
Veröffentlicht: (2025)
LakeVisage: Towards Scalable, Flexible and Interactive Visualization Recommendation for Data Discovery over Data Lakes
von: Hu, Yihao, et al.
Veröffentlicht: (2025)
von: Hu, Yihao, et al.
Veröffentlicht: (2025)
Blueprinting the Cloud: Unifying and Automatically Optimizing Cloud Data Infrastructures with BRAD -- Extended Version
von: Yu, Geoffrey X., et al.
Veröffentlicht: (2024)
von: Yu, Geoffrey X., et al.
Veröffentlicht: (2024)
Efficient Data Ingestion in Cloud-based architecture: a Data Engineering Design Pattern Proposal
von: Rucco, Chiara, et al.
Veröffentlicht: (2025)
von: Rucco, Chiara, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Privacy-Enhanced Database Synthesis for Benchmark Publishing (Technical Report)
von: Ge, Yunqing, et al.
Veröffentlicht: (2024) -
Ultraverse: A System-Centric Framework for Efficient What-If Analysis for Database-Intensive Web Applications
von: Ko, Ronny, et al.
Veröffentlicht: (2022) -
ShapleyPipe: Hierarchical Shapley Search for Data Preparation Pipeline Construction
von: Chang, Jing, et al.
Veröffentlicht: (2025) -
LLaPipe: LLM-Guided Reinforcement Learning for Automated Data Preparation Pipeline Construction
von: Chang, Jing, et al.
Veröffentlicht: (2025) -
SoftPipe: A Soft-Guided Reinforcement Learning Framework for Automated Data Preparation
von: Chang, Jing, et al.
Veröffentlicht: (2025)