A Survey of Pipeline Tools for Data Engineering
Fuente:
arXiv
Salvato in:
| Autori principali: | Mbata, Anthony, Sripada, Yaji, Zhong, Mingjun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluation of Pipelines for Data Integration into Knowledge Graphs
di: Hofer, Marvin, et al.
Pubblicazione: (2026)
di: Hofer, Marvin, et al.
Pubblicazione: (2026)
KGpipe: Generation and Evaluation of Pipelines for Data Integration into Knowledge Graphs
di: Hofer, Marvin, et al.
Pubblicazione: (2025)
di: Hofer, Marvin, et al.
Pubblicazione: (2025)
Trajectory Data Management and Mining: A Survey from Deep Learning to the LLM Era
di: Chen, Wei, et al.
Pubblicazione: (2024)
di: Chen, Wei, et al.
Pubblicazione: (2024)
Graph Learning in the Era of LLMs: A Survey from the Perspective of Data, Models, and Tasks
di: Li, Xunkai, et al.
Pubblicazione: (2024)
di: Li, Xunkai, et al.
Pubblicazione: (2024)
Modyn: Data-Centric Machine Learning Pipeline Orchestration
di: Böther, Maximilian, et al.
Pubblicazione: (2023)
di: Böther, Maximilian, et al.
Pubblicazione: (2023)
Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs
di: Zhou, Wei, et al.
Pubblicazione: (2026)
di: Zhou, Wei, et al.
Pubblicazione: (2026)
When Large Language Models Meet Vector Databases: A Survey
di: Jing, Zhi, et al.
Pubblicazione: (2024)
di: Jing, Zhi, et al.
Pubblicazione: (2024)
Data Agent: A Holistic Architecture for Orchestrating Data+AI Ecosystems
di: Sun, Zhaoyan, et al.
Pubblicazione: (2025)
di: Sun, Zhaoyan, et al.
Pubblicazione: (2025)
A Survey on Time-Series Distance Measures
di: Paparrizos, John, et al.
Pubblicazione: (2024)
di: Paparrizos, John, et al.
Pubblicazione: (2024)
Jellyfish: A Large Language Model for Data Preprocessing
di: Zhang, Haochen, et al.
Pubblicazione: (2023)
di: Zhang, Haochen, et al.
Pubblicazione: (2023)
LLM-Enhanced Data Management
di: Zhou, Xuanhe, et al.
Pubblicazione: (2024)
di: Zhou, Xuanhe, et al.
Pubblicazione: (2024)
Database Entity Recognition with Data Augmentation and Deep Learning
di: Fu, Zikun, et al.
Pubblicazione: (2025)
di: Fu, Zikun, et al.
Pubblicazione: (2025)
A Survey of LLM $\times$ DATA
di: Zhou, Xuanhe, et al.
Pubblicazione: (2025)
di: Zhou, Xuanhe, et al.
Pubblicazione: (2025)
QUIS: Question-guided Insights Generation for Automated Exploratory Data Analysis
di: Manatkar, Abhijit, et al.
Pubblicazione: (2024)
di: Manatkar, Abhijit, et al.
Pubblicazione: (2024)
Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models
di: Liu, Wei, et al.
Pubblicazione: (2026)
di: Liu, Wei, et al.
Pubblicazione: (2026)
PMKLC: Parallel Multi-Knowledge Learning-based Lossless Compression for Large-Scale Genomics Database
di: Sun, Hui, et al.
Pubblicazione: (2025)
di: Sun, Hui, et al.
Pubblicazione: (2025)
SQL-GEN: Bridging the Dialect Gap for Text-to-SQL Via Synthetic Data And Model Merging
di: Pourreza, Mohammadreza, et al.
Pubblicazione: (2024)
di: Pourreza, Mohammadreza, et al.
Pubblicazione: (2024)
Mixtera: A Data Plane for Foundation Model Training
di: Böther, Maximilian, et al.
Pubblicazione: (2025)
di: Böther, Maximilian, et al.
Pubblicazione: (2025)
Mind the Data Gap: Bridging LLMs to Enterprise Data Integration
di: Kayali, Moe, et al.
Pubblicazione: (2024)
di: Kayali, Moe, et al.
Pubblicazione: (2024)
Intelligent Cross-Organizational Process Mining: A Survey and New Perspectives
di: Yang, Yiyuan, et al.
Pubblicazione: (2024)
di: Yang, Yiyuan, et al.
Pubblicazione: (2024)
Cortex AISQL: A Production SQL Engine for Unstructured Data
di: Liskowski, Paweł, et al.
Pubblicazione: (2025)
di: Liskowski, Paweł, et al.
Pubblicazione: (2025)
Synthetic Tabular Data Detection In the Wild
di: Kindji, G. Charbel N., et al.
Pubblicazione: (2025)
di: Kindji, G. Charbel N., et al.
Pubblicazione: (2025)
Towards Controllable Time Series Generation
di: Bao, Yifan, et al.
Pubblicazione: (2024)
di: Bao, Yifan, et al.
Pubblicazione: (2024)
TableGPT2: A Large Multimodal Model with Tabular Data Integration
di: Su, Aofeng, et al.
Pubblicazione: (2024)
di: Su, Aofeng, et al.
Pubblicazione: (2024)
MontePrep: Monte-Carlo-Driven Automatic Data Preparation without Target Data Instances
di: Ge, Congcong, et al.
Pubblicazione: (2025)
di: Ge, Congcong, et al.
Pubblicazione: (2025)
TabSketchFM: Sketch-based Tabular Representation Learning for Data Discovery over Data Lakes
di: Khatiwada, Aamod, et al.
Pubblicazione: (2024)
di: Khatiwada, Aamod, et al.
Pubblicazione: (2024)
Generating the Traces You Need: A Conditional Generative Model for Process Mining Data
di: Graziosi, Riccardo, et al.
Pubblicazione: (2024)
di: Graziosi, Riccardo, et al.
Pubblicazione: (2024)
Integrating Meteorological and Operational Data: A Novel Approach to Understanding Railway Delays in Finland
di: Borin, Vinicius Pozzobon, et al.
Pubblicazione: (2026)
di: Borin, Vinicius Pozzobon, et al.
Pubblicazione: (2026)
Data Science: a Natural Ecosystem
di: Porcu, Emilio, et al.
Pubblicazione: (2025)
di: Porcu, Emilio, et al.
Pubblicazione: (2025)
OpenMLDB: A Real-Time Relational Data Feature Computation System for Online ML
di: Zhou, Xuanhe, et al.
Pubblicazione: (2025)
di: Zhou, Xuanhe, et al.
Pubblicazione: (2025)
Data Collection and Labeling Techniques for Machine Learning
di: Huang, Qianyu, et al.
Pubblicazione: (2024)
di: Huang, Qianyu, et al.
Pubblicazione: (2024)
Making LLMs Work for Enterprise Data Tasks
di: Demiralp, Çağatay, et al.
Pubblicazione: (2024)
di: Demiralp, Çağatay, et al.
Pubblicazione: (2024)
Datum-wise Transformer for Synthetic Tabular Data Detection in the Wild
di: Kindji, G. Charbel N., et al.
Pubblicazione: (2025)
di: Kindji, G. Charbel N., et al.
Pubblicazione: (2025)
GenDB: The Next Generation of Query Processing -- Synthesized, Not Engineered
di: Lao, Jiale, et al.
Pubblicazione: (2026)
di: Lao, Jiale, et al.
Pubblicazione: (2026)
An Automated LLM-based Pipeline for Asset-Level Database Creation to Assess Deforestation Impact
di: Menon, Avanija, et al.
Pubblicazione: (2025)
di: Menon, Avanija, et al.
Pubblicazione: (2025)
NeurDB: An AI-powered Autonomous Data System
di: Ooi, Beng Chin, et al.
Pubblicazione: (2024)
di: Ooi, Beng Chin, et al.
Pubblicazione: (2024)
Forgetting by Pruning: Data Deletion in Join Cardinality Estimation
di: He, Chaowei, et al.
Pubblicazione: (2025)
di: He, Chaowei, et al.
Pubblicazione: (2025)
Process Mining for Unstructured Data: Challenges and Research Directions
di: Koschmider, Agnes, et al.
Pubblicazione: (2023)
di: Koschmider, Agnes, et al.
Pubblicazione: (2023)
KDSelector: A Knowledge-Enhanced and Data-Efficient Model Selector Learning Framework for Time Series Anomaly Detection
di: Liang, Zhiyu, et al.
Pubblicazione: (2025)
di: Liang, Zhiyu, et al.
Pubblicazione: (2025)
LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning
di: Lin, Xiaotian, et al.
Pubblicazione: (2025)
di: Lin, Xiaotian, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Evaluation of Pipelines for Data Integration into Knowledge Graphs
di: Hofer, Marvin, et al.
Pubblicazione: (2026) -
KGpipe: Generation and Evaluation of Pipelines for Data Integration into Knowledge Graphs
di: Hofer, Marvin, et al.
Pubblicazione: (2025) -
Trajectory Data Management and Mining: A Survey from Deep Learning to the LLM Era
di: Chen, Wei, et al.
Pubblicazione: (2024) -
Graph Learning in the Era of LLMs: A Survey from the Perspective of Data, Models, and Tasks
di: Li, Xunkai, et al.
Pubblicazione: (2024) -
Modyn: Data-Centric Machine Learning Pipeline Orchestration
di: Böther, Maximilian, et al.
Pubblicazione: (2023)