DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Hao, Ma, Xiaochen, Liu, Zhou, Wong, Zhen Hao, Zhao, Zhengyang, Meng, Zimo, He, Runming, Shen, Chengyu, Cai, Qifeng, Han, Zhaoyang, Qiang, Meiyi, Feng, Yalin, Bai, Tianyi, Pan, Zewei, Guo, Ziyi, Jiang, Yizhen, Deng, Jingwen, You, Qijie, Lai, Peichao, Guo, Tianyu, Tsai, Chi Hsu, Feng, Hengyi, Hu, Rui, Yu, Wenkai, Niu, Junbo, Zeng, Bohan, An, Ruichuan, Ma, Lu, Huang, Jihao, Zheng, Yaowei, He, Conghui, Tang, Linpeng, Cui, Bin, E, Weinan, Zhang, Wentao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models
by: Liang, Hao, et al.
Published: (2026)
by: Liang, Hao, et al.
Published: (2026)
Towards Next-Generation LLM Training: From the Data-Centric Perspective
by: Liang, Hao, et al.
Published: (2026)
by: Liang, Hao, et al.
Published: (2026)
ANDES: Agent Native Data Evolving Synthesis Tool for Autonomous Instruction Alignment
by: Zhao, Zhengyang, et al.
Published: (2026)
by: Zhao, Zhengyang, et al.
Published: (2026)
One-Eval: An Agentic System for Automated and Traceable LLM Evaluation
by: Shen, Chengyu, et al.
Published: (2026)
by: Shen, Chengyu, et al.
Published: (2026)
Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
by: Ma, Lu, et al.
Published: (2025)
by: Ma, Lu, et al.
Published: (2025)
Causify DataFlow: A Framework For High-performance Machine Learning Stream Computing
by: Saggese, Giacinto Paolo, et al.
Published: (2025)
by: Saggese, Giacinto Paolo, et al.
Published: (2025)
FlipVQA: Scaling Multi-modal Instruction Tuning via Textbook-to-Knowledge Synthesis
by: Wong, Zhen Hao, et al.
Published: (2025)
by: Wong, Zhen Hao, et al.
Published: (2025)
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos
by: Feng, Hengyi, et al.
Published: (2026)
by: Feng, Hengyi, et al.
Published: (2026)
MathClean: A Benchmark for Synthetic Mathematical Data Cleaning
by: Liang, Hao, et al.
Published: (2025)
by: Liang, Hao, et al.
Published: (2025)
Let's Verify Math Questions Step by Step
by: Shen, Chengyu, et al.
Published: (2025)
by: Shen, Chengyu, et al.
Published: (2025)
LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
by: Cai, Qifeng, et al.
Published: (2025)
by: Cai, Qifeng, et al.
Published: (2025)
Generative Giants, Retrieval Weaklings: Why do Multimodal Large Language Models Fail at Multimodal Retrieval?
by: Feng, Hengyi, et al.
Published: (2025)
by: Feng, Hengyi, et al.
Published: (2025)
Topic Over Source: The Key to Effective Data Mixing for Language Models Pre-training
by: Peng, Jiahui, et al.
Published: (2025)
by: Peng, Jiahui, et al.
Published: (2025)
Enhancing Crash Frequency Modeling Based on Augmented Multi-Type Data by Hybrid VAE-Diffusion-Based Generative Neural Networks
by: Chen, Junlan, et al.
Published: (2025)
by: Chen, Junlan, et al.
Published: (2025)
When grandparents step back: Fertility intentions and policy responses amid delayed retirement
by: Chen, He, et al.
Published: (2025)
by: Chen, He, et al.
Published: (2025)
Strong and weak well-posedness of McKean-Vlasov SDEs driven by $α$-stable processes under unified condition
by: Hao, Zimo
Published: (2025)
by: Hao, Zimo
Published: (2025)
Synth-Empathy: Towards High-Quality Synthetic Empathy Data
by: Liang, Hao, et al.
Published: (2024)
by: Liang, Hao, et al.
Published: (2024)
Closing the Data Loop: Using OpenDataArena to Engineer Superior Training Datasets
by: Gao, Xin, et al.
Published: (2025)
by: Gao, Xin, et al.
Published: (2025)
Spectral Property-Driven Data Augmentation for Hyperspectral Single-Source Domain Generalization
by: Chen, Taiqin, et al.
Published: (2026)
by: Chen, Taiqin, et al.
Published: (2026)
LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning
by: Wong, Zhen Hao, et al.
Published: (2025)
by: Wong, Zhen Hao, et al.
Published: (2025)
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
by: Liang, Hao, et al.
Published: (2026)
by: Liang, Hao, et al.
Published: (2026)
PUAL: A Classifier on Trifurcate Positive-Unlabeled Data
by: Wang, Xiaoke, et al.
Published: (2024)
by: Wang, Xiaoke, et al.
Published: (2024)
Deep Feature Embedding for Tabular Data
by: Wu, Yuqian, et al.
Published: (2024)
by: Wu, Yuqian, et al.
Published: (2024)
LIAS Promotes Cuproptosis in Prostate Cancer Cells by Suppressing Glycolysis via the p53 Signaling Pathway
by: Zhe Tang, et al.
Published: (2025)
by: Zhe Tang, et al.
Published: (2025)
GIFT: Reconciling Post-Training Objectives via Finite-Temperature Gibbs Initialization
by: Zhao, Zhengyang, et al.
Published: (2026)
by: Zhao, Zhengyang, et al.
Published: (2026)
On the Impact of Uncertainty and Calibration on Likelihood-Ratio Membership Inference Attacks
by: Zhu, Meiyi, et al.
Published: (2024)
by: Zhu, Meiyi, et al.
Published: (2024)
Attention-Based Feature Online Conformal Prediction for Time Series
by: Zhu, Meiyi, et al.
Published: (2025)
by: Zhu, Meiyi, et al.
Published: (2025)
Towards Reliable Neural Optimizers: Permutation-Equivariant Neural Approximation in Dynamic Data Driven Applications Systems
by: Li, Meiyi, et al.
Published: (2025)
by: Li, Meiyi, et al.
Published: (2025)
AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG
by: You, Qijie, et al.
Published: (2026)
by: You, Qijie, et al.
Published: (2026)
Beyond Normality: Reliable A/B Testing with Non-Gaussian Data
by: Gong, Junpeng, et al.
Published: (2025)
by: Gong, Junpeng, et al.
Published: (2025)
A Deep Learning Approach to Anomaly Detection in High-Frequency Trading Data
by: Bao, Qiuliuyang, et al.
Published: (2025)
by: Bao, Qiuliuyang, et al.
Published: (2025)
IoDResearch: Deep Research on Private Heterogeneous Data via the Internet of Data
by: Shi, Zhuofan, et al.
Published: (2025)
by: Shi, Zhuofan, et al.
Published: (2025)
Language Modeling on Tabular Data: A Survey of Foundations, Techniques and Evolution
by: Ruan, Yucheng, et al.
Published: (2024)
by: Ruan, Yucheng, et al.
Published: (2024)
DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception
by: Zhao, Zhiyuan, et al.
Published: (2024)
by: Zhao, Zhiyuan, et al.
Published: (2024)
Quantitative bounds for critically bounded solutions to the three-dimensional Navier-Stokes equations in Lorentz spaces
by: Feng, Wen, et al.
Published: (2022)
by: Feng, Wen, et al.
Published: (2022)
Robust Training for Speaker Verification against Noisy Labels
by: Fang, Zhihua, et al.
Published: (2022)
by: Fang, Zhihua, et al.
Published: (2022)
Probable Event Constrained Optimization and A Data-embedded Solution Paradigm
by: Li, Qifeng
Published: (2022)
by: Li, Qifeng
Published: (2022)
Can Modifying Data Address Graph Domain Adaptation?
by: Huang, Renhong, et al.
Published: (2024)
by: Huang, Renhong, et al.
Published: (2024)
Changes of filial responsibility norms under public long‐term care insurance in China
by: Qifeng Ma, et al.
Published: (2025)
by: Qifeng Ma, et al.
Published: (2025)
Safe Data-Driven Control and Dynamical Learning via Constrained Neural Architectures and Koopman Operators
by: Feng, Lin, et al.
Published: (2026)
by: Feng, Lin, et al.
Published: (2026)
Similar Items
-
DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models
by: Liang, Hao, et al.
Published: (2026) -
Towards Next-Generation LLM Training: From the Data-Centric Perspective
by: Liang, Hao, et al.
Published: (2026) -
ANDES: Agent Native Data Evolving Synthesis Tool for Autonomous Instruction Alignment
by: Zhao, Zhengyang, et al.
Published: (2026) -
One-Eval: An Agentic System for Automated and Traceable LLM Evaluation
by: Shen, Chengyu, et al.
Published: (2026) -
Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
by: Ma, Lu, et al.
Published: (2025)