Mixtera: A Data Plane for Foundation Model Training
Fuente:
arXiv
Salvato in:
| Autori principali: | Böther, Maximilian, Yao, Xiaozhe, Kerimoglu, Tolga, Graur, Dan, Gsteiger, Viktor, Klimovic, Ana |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Modyn: Data-Centric Machine Learning Pipeline Orchestration
di: Böther, Maximilian, et al.
Pubblicazione: (2023)
di: Böther, Maximilian, et al.
Pubblicazione: (2023)
No Need to Train Your RDB Foundation Model
di: Xu, Linjie, et al.
Pubblicazione: (2026)
di: Xu, Linjie, et al.
Pubblicazione: (2026)
Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data
di: Ranjan, Rishabh, et al.
Pubblicazione: (2025)
di: Ranjan, Rishabh, et al.
Pubblicazione: (2025)
Position: Foundation Models for Tabular Data within Systemic Contexts Need Grounding
di: Klein, Tassilo, et al.
Pubblicazione: (2025)
di: Klein, Tassilo, et al.
Pubblicazione: (2025)
PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models
di: Kothapalli, Vignesh, et al.
Pubblicazione: (2026)
di: Kothapalli, Vignesh, et al.
Pubblicazione: (2026)
TensorBank: Tensor Lakehouse for Foundation Model Training
di: Kienzler, Romeo, et al.
Pubblicazione: (2023)
di: Kienzler, Romeo, et al.
Pubblicazione: (2023)
On Distributed Larger-Than-Memory Subset Selection With Pairwise Submodular Functions
di: Böther, Maximilian, et al.
Pubblicazione: (2024)
di: Böther, Maximilian, et al.
Pubblicazione: (2024)
Tabular Foundation Models Can Learn Association Rules
di: Karabulut, Erkan, et al.
Pubblicazione: (2026)
di: Karabulut, Erkan, et al.
Pubblicazione: (2026)
Efficient Data Access Paths for Mixed Vector-Relational Search
di: Sanca, Viktor, et al.
Pubblicazione: (2024)
di: Sanca, Viktor, et al.
Pubblicazione: (2024)
Griffin: Towards a Graph-Centric Relational Database Foundation Model
di: Wang, Yanbo, et al.
Pubblicazione: (2025)
di: Wang, Yanbo, et al.
Pubblicazione: (2025)
HoneyBee: A Scalable Modular Framework for Creating Multimodal Oncology Datasets with Foundational Embedding Models
di: Tripathi, Aakash, et al.
Pubblicazione: (2024)
di: Tripathi, Aakash, et al.
Pubblicazione: (2024)
Optimizing Context-Enhanced Relational Joins
di: Sanca, Viktor, et al.
Pubblicazione: (2023)
di: Sanca, Viktor, et al.
Pubblicazione: (2023)
Process Mining for Unstructured Data: Challenges and Research Directions
di: Koschmider, Agnes, et al.
Pubblicazione: (2023)
di: Koschmider, Agnes, et al.
Pubblicazione: (2023)
Integrating Meteorological and Operational Data: A Novel Approach to Understanding Railway Delays in Finland
di: Borin, Vinicius Pozzobon, et al.
Pubblicazione: (2026)
di: Borin, Vinicius Pozzobon, et al.
Pubblicazione: (2026)
Relational Deep Learning: Challenges, Foundations and Next-Generation Architectures
di: Dwivedi, Vijay Prakash, et al.
Pubblicazione: (2025)
di: Dwivedi, Vijay Prakash, et al.
Pubblicazione: (2025)
TableGPT2: A Large Multimodal Model with Tabular Data Integration
di: Su, Aofeng, et al.
Pubblicazione: (2024)
di: Su, Aofeng, et al.
Pubblicazione: (2024)
Generating the Traces You Need: A Conditional Generative Model for Process Mining Data
di: Graziosi, Riccardo, et al.
Pubblicazione: (2024)
di: Graziosi, Riccardo, et al.
Pubblicazione: (2024)
DiffImpute: Tabular Data Imputation With Denoising Diffusion Probabilistic Model
di: Wen, Yizhu, et al.
Pubblicazione: (2024)
di: Wen, Yizhu, et al.
Pubblicazione: (2024)
KDSelector: A Knowledge-Enhanced and Data-Efficient Model Selector Learning Framework for Time Series Anomaly Detection
di: Liang, Zhiyu, et al.
Pubblicazione: (2025)
di: Liang, Zhiyu, et al.
Pubblicazione: (2025)
Towards Privacy-Preserving Relational Data Synthesis via Probabilistic Relational Models
di: Luttermann, Malte, et al.
Pubblicazione: (2024)
di: Luttermann, Malte, et al.
Pubblicazione: (2024)
Accelerating Storage-Based Training for Graph Neural Networks
di: Jang, Myung-Hwan, et al.
Pubblicazione: (2026)
di: Jang, Myung-Hwan, et al.
Pubblicazione: (2026)
Efficient Training-Free Online Routing for High-Volume Multi-LLM Serving
di: Wu, Fangzhou, et al.
Pubblicazione: (2025)
di: Wu, Fangzhou, et al.
Pubblicazione: (2025)
Mind the Data Gap: Bridging LLMs to Enterprise Data Integration
di: Kayali, Moe, et al.
Pubblicazione: (2024)
di: Kayali, Moe, et al.
Pubblicazione: (2024)
Cortex AISQL: A Production SQL Engine for Unstructured Data
di: Liskowski, Paweł, et al.
Pubblicazione: (2025)
di: Liskowski, Paweł, et al.
Pubblicazione: (2025)
MontePrep: Monte-Carlo-Driven Automatic Data Preparation without Target Data Instances
di: Ge, Congcong, et al.
Pubblicazione: (2025)
di: Ge, Congcong, et al.
Pubblicazione: (2025)
TabSketchFM: Sketch-based Tabular Representation Learning for Data Discovery over Data Lakes
di: Khatiwada, Aamod, et al.
Pubblicazione: (2024)
di: Khatiwada, Aamod, et al.
Pubblicazione: (2024)
Data Science: a Natural Ecosystem
di: Porcu, Emilio, et al.
Pubblicazione: (2025)
di: Porcu, Emilio, et al.
Pubblicazione: (2025)
OpenMLDB: A Real-Time Relational Data Feature Computation System for Online ML
di: Zhou, Xuanhe, et al.
Pubblicazione: (2025)
di: Zhou, Xuanhe, et al.
Pubblicazione: (2025)
Evaluation of Pipelines for Data Integration into Knowledge Graphs
di: Hofer, Marvin, et al.
Pubblicazione: (2026)
di: Hofer, Marvin, et al.
Pubblicazione: (2026)
Data Collection and Labeling Techniques for Machine Learning
di: Huang, Qianyu, et al.
Pubblicazione: (2024)
di: Huang, Qianyu, et al.
Pubblicazione: (2024)
Making LLMs Work for Enterprise Data Tasks
di: Demiralp, Çağatay, et al.
Pubblicazione: (2024)
di: Demiralp, Çağatay, et al.
Pubblicazione: (2024)
ExCoT: Optimizing Reasoning for Text-to-SQL with Execution Feedback
di: Zhai, Bohan, et al.
Pubblicazione: (2025)
di: Zhai, Bohan, et al.
Pubblicazione: (2025)
Forgetting by Pruning: Data Deletion in Join Cardinality Estimation
di: He, Chaowei, et al.
Pubblicazione: (2025)
di: He, Chaowei, et al.
Pubblicazione: (2025)
NeurDB: An AI-powered Autonomous Data System
di: Ooi, Beng Chin, et al.
Pubblicazione: (2024)
di: Ooi, Beng Chin, et al.
Pubblicazione: (2024)
A Lightweight Learned Cardinality Estimation Model
di: Zhu, Yaoyu, et al.
Pubblicazione: (2025)
di: Zhu, Yaoyu, et al.
Pubblicazione: (2025)
LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning
di: Lin, Xiaotian, et al.
Pubblicazione: (2025)
di: Lin, Xiaotian, et al.
Pubblicazione: (2025)
Towards Synthesizing High-Dimensional Tabular Data with Limited Samples
di: Li, Zuqing, et al.
Pubblicazione: (2025)
di: Li, Zuqing, et al.
Pubblicazione: (2025)
KGpipe: Generation and Evaluation of Pipelines for Data Integration into Knowledge Graphs
di: Hofer, Marvin, et al.
Pubblicazione: (2025)
di: Hofer, Marvin, et al.
Pubblicazione: (2025)
Robust Detection of Synthetic Tabular Data under Schema Variability
di: Kindji, G. Charbel N., et al.
Pubblicazione: (2025)
di: Kindji, G. Charbel N., et al.
Pubblicazione: (2025)
ILAEDA: An Imitation Learning Based Approach for Automatic Exploratory Data Analysis
di: Manatkar, Abhijit, et al.
Pubblicazione: (2024)
di: Manatkar, Abhijit, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Modyn: Data-Centric Machine Learning Pipeline Orchestration
di: Böther, Maximilian, et al.
Pubblicazione: (2023) -
No Need to Train Your RDB Foundation Model
di: Xu, Linjie, et al.
Pubblicazione: (2026) -
Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data
di: Ranjan, Rishabh, et al.
Pubblicazione: (2025) -
Position: Foundation Models for Tabular Data within Systemic Contexts Need Grounding
di: Klein, Tassilo, et al.
Pubblicazione: (2025) -
PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models
di: Kothapalli, Vignesh, et al.
Pubblicazione: (2026)