Automated Planning for Optimal Data Pipeline Instantiation
Fuente:
arXiv
Saved in:
| Main Authors: | Amado, Leonardo Rosa, Vogel, Adriano, Griebler, Dalvan, Licks, Gabriel Paludo, Simon, Eric, Meneguzzi, Felipe |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LOG.io: Unified Rollback Recovery and Data Lineage Capture for Distributed Data Pipelines
by: Simon, Eric, et al.
Published: (2025)
by: Simon, Eric, et al.
Published: (2025)
NPB-Rust: NAS Parallel Benchmarks in Rust
by: Martins, Eduardo M., et al.
Published: (2025)
by: Martins, Eduardo M., et al.
Published: (2025)
WORKSWORLD: A Domain for Integrated Numeric Planning and Scheduling of Distributed Pipelined Workflows
by: Paul, Taylor, et al.
Published: (2026)
by: Paul, Taylor, et al.
Published: (2026)
Artificial Intelligence for Cost-Aware Resource Prediction in Big Data Pipelines
by: Goyal, Harshit
Published: (2025)
by: Goyal, Harshit
Published: (2025)
AdaPtis: Reducing Pipeline Bubbles with Adaptive Pipeline Parallelism on Heterogeneous Models
by: Guo, Jihu, et al.
Published: (2025)
by: Guo, Jihu, et al.
Published: (2025)
FreeRide: Harvesting Bubbles in Pipeline Parallelism
by: Zhang, Jiashu, et al.
Published: (2024)
by: Zhang, Jiashu, et al.
Published: (2024)
Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
by: Wagenländer, Marcel, et al.
Published: (2026)
by: Wagenländer, Marcel, et al.
Published: (2026)
TimelyFreeze: Adaptive Parameter Freezing Mechanism for Pipeline Parallelism
by: Cho, Seonghye, et al.
Published: (2026)
by: Cho, Seonghye, et al.
Published: (2026)
Understand and Accelerate Memory Processing Pipeline for Large Language Model Inference
by: He, Zifan, et al.
Published: (2026)
by: He, Zifan, et al.
Published: (2026)
DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline
by: Xue, Zhenliang, et al.
Published: (2025)
by: Xue, Zhenliang, et al.
Published: (2025)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
by: Liu, Xing, et al.
Published: (2025)
by: Liu, Xing, et al.
Published: (2025)
Building a Correct-by-Design Lakehouse. Data Contracts, Versioning, and Transactional Pipelines for Humans and Agents
by: Sheng, Weiming, et al.
Published: (2026)
by: Sheng, Weiming, et al.
Published: (2026)
InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training
by: Wang, Shiju, et al.
Published: (2025)
by: Wang, Shiju, et al.
Published: (2025)
Transforming Future Data Center Operations and Management via Physical AI
by: Cao, Zhiwei, et al.
Published: (2025)
by: Cao, Zhiwei, et al.
Published: (2025)
Automated Road Safety: Enhancing Sign and Surface Damage Detection with AI
by: Merolla, Davide, et al.
Published: (2024)
by: Merolla, Davide, et al.
Published: (2024)
Capacity Planning and Scheduling for Jobs with Uncertainty in Resource Usage and Duration
by: Patra, Sunandita, et al.
Published: (2025)
by: Patra, Sunandita, et al.
Published: (2025)
Meeting SLOs, Slashing Hours: Automated Enterprise LLM Optimization with OptiKIT
by: Santavas, Nicholas, et al.
Published: (2026)
by: Santavas, Nicholas, et al.
Published: (2026)
Connecting Large Language Models with Blockchain: Advancing the Evolution of Smart Contracts from Automation to Intelligence
by: Xian, Youquan, et al.
Published: (2024)
by: Xian, Youquan, et al.
Published: (2024)
Automated Database Indexing using Model-free Reinforcement Learning
by: Licks, Gabriel Paludo, et al.
Published: (2020)
by: Licks, Gabriel Paludo, et al.
Published: (2020)
Zero Bubble Pipeline Parallelism
by: Qi, Penghui, et al.
Published: (2023)
by: Qi, Penghui, et al.
Published: (2023)
Data-Juicer 2.0: Cloud-Scale Adaptive Data Processing for and with Foundation Models
by: Chen, Daoyuan, et al.
Published: (2024)
by: Chen, Daoyuan, et al.
Published: (2024)
Modyn: Data-Centric Machine Learning Pipeline Orchestration
by: Böther, Maximilian, et al.
Published: (2023)
by: Böther, Maximilian, et al.
Published: (2023)
DataCenterGym: A Physics-Grounded Simulator for Multi-Objective Data Center Scheduling
by: Pathak, Nilavra, et al.
Published: (2026)
by: Pathak, Nilavra, et al.
Published: (2026)
Viability and Performance of a Private LLM Server for SMBs: A Benchmark Analysis of Qwen3-30B on Consumer-Grade Hardware
by: Khalil, Alex, et al.
Published: (2025)
by: Khalil, Alex, et al.
Published: (2025)
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
by: Kamahori, Keisuke, et al.
Published: (2026)
by: Kamahori, Keisuke, et al.
Published: (2026)
Performance Evaluation of LLMs in Automated RDF Knowledge Graph Generation
by: Martin, Ioana Ramona, et al.
Published: (2026)
by: Martin, Ioana Ramona, et al.
Published: (2026)
Isambard-AI: a leadership class supercomputer optimised specifically for Artificial Intelligence
by: McIntosh-Smith, Simon, et al.
Published: (2024)
by: McIntosh-Smith, Simon, et al.
Published: (2024)
Full Scaling Automation for Sustainable Development of Green Data Centers
by: Wang, Shiyu, et al.
Published: (2023)
by: Wang, Shiyu, et al.
Published: (2023)
Scaling Performance of Large Language Model Pretraining
by: Interrante-Grant, Alexander, et al.
Published: (2025)
by: Interrante-Grant, Alexander, et al.
Published: (2025)
ALPACA -- Adaptive Learning Pipeline for Comprehensive AI
by: Torka, Simon, et al.
Published: (2024)
by: Torka, Simon, et al.
Published: (2024)
LLM as HPC Expert: Extending RAG Architecture for HPC Data
by: Miyashita, Yusuke, et al.
Published: (2024)
by: Miyashita, Yusuke, et al.
Published: (2024)
Scalable Cloud-Native Architectures for Intelligent PMU Data Processing
by: Chockalingam, Nachiappan, et al.
Published: (2025)
by: Chockalingam, Nachiappan, et al.
Published: (2025)
AgileLog: A Forkable Shared Log for Agents on Data Streams
by: Bhat, Shreesha G., et al.
Published: (2026)
by: Bhat, Shreesha G., et al.
Published: (2026)
SimpleFSDP: Simpler Fully Sharded Data Parallel with torch.compile
by: Zhang, Ruisi, et al.
Published: (2024)
by: Zhang, Ruisi, et al.
Published: (2024)
Towards using Reinforcement Learning for Scaling and Data Replication in Cloud Systems
by: Mokadem, Riad, et al.
Published: (2024)
by: Mokadem, Riad, et al.
Published: (2024)
Benchmarking of CPU-intensive Stream Data Processing in The Edge Computing Systems
by: Szydlo, Tomasz, et al.
Published: (2025)
by: Szydlo, Tomasz, et al.
Published: (2025)
Ensemble Method for System Failure Detection Using Large-Scale Telemetry Data
by: Mudgal, Priyanka, et al.
Published: (2024)
by: Mudgal, Priyanka, et al.
Published: (2024)
MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training
by: Zhao, Juntao, et al.
Published: (2025)
by: Zhao, Juntao, et al.
Published: (2025)
Reinforcement Learning-driven Data-intensive Workflow Scheduling for Volunteer Edge-Cloud
by: Mounesan, Motahare, et al.
Published: (2024)
by: Mounesan, Motahare, et al.
Published: (2024)
Training Through Failure: Effects of Data Consistency in Parallel Machine Learning Training
by: Cao, Ray, et al.
Published: (2024)
by: Cao, Ray, et al.
Published: (2024)
Similar Items
-
LOG.io: Unified Rollback Recovery and Data Lineage Capture for Distributed Data Pipelines
by: Simon, Eric, et al.
Published: (2025) -
NPB-Rust: NAS Parallel Benchmarks in Rust
by: Martins, Eduardo M., et al.
Published: (2025) -
WORKSWORLD: A Domain for Integrated Numeric Planning and Scheduling of Distributed Pipelined Workflows
by: Paul, Taylor, et al.
Published: (2026) -
Artificial Intelligence for Cost-Aware Resource Prediction in Big Data Pipelines
by: Goyal, Harshit
Published: (2025) -
AdaPtis: Reducing Pipeline Bubbles with Adaptive Pipeline Parallelism on Heterogeneous Models
by: Guo, Jihu, et al.
Published: (2025)