DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Hao, Zhao, Zhengyang, Qiang, Meiyi, Chen, Mingrui, Ma, Lu, Yu, Rongyi, Feng, Hengyi, Sun, Shixuan, Meng, Zimo, Ma, Xiaochen, Yang, Xuanlin, Cai, Qifeng, An, Ruichuan, Zeng, Bohan, Wong, Zhen Hao, Shen, Chengyu, He, Runming, Han, Zhaoyang, Zheng, Yaowei, Fu, Fangcheng, He, Conghui, Cui, Bin, Li, Zhiyu, E, Weinan, Zhang, Wentao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
von: Liang, Hao, et al.
Veröffentlicht: (2025)
von: Liang, Hao, et al.
Veröffentlicht: (2025)
Towards Next-Generation LLM Training: From the Data-Centric Perspective
von: Liang, Hao, et al.
Veröffentlicht: (2026)
von: Liang, Hao, et al.
Veröffentlicht: (2026)
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos
von: Feng, Hengyi, et al.
Veröffentlicht: (2026)
von: Feng, Hengyi, et al.
Veröffentlicht: (2026)
Let's Verify Math Questions Step by Step
von: Shen, Chengyu, et al.
Veröffentlicht: (2025)
von: Shen, Chengyu, et al.
Veröffentlicht: (2025)
Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
von: Ma, Lu, et al.
Veröffentlicht: (2025)
von: Ma, Lu, et al.
Veröffentlicht: (2025)
LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
von: Cai, Qifeng, et al.
Veröffentlicht: (2025)
von: Cai, Qifeng, et al.
Veröffentlicht: (2025)
ANDES: Agent Native Data Evolving Synthesis Tool for Autonomous Instruction Alignment
von: Zhao, Zhengyang, et al.
Veröffentlicht: (2026)
von: Zhao, Zhengyang, et al.
Veröffentlicht: (2026)
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
von: Liang, Hao, et al.
Veröffentlicht: (2026)
von: Liang, Hao, et al.
Veröffentlicht: (2026)
GIFT: Reconciling Post-Training Objectives via Finite-Temperature Gibbs Initialization
von: Zhao, Zhengyang, et al.
Veröffentlicht: (2026)
von: Zhao, Zhengyang, et al.
Veröffentlicht: (2026)
One-Eval: An Agentic System for Automated and Traceable LLM Evaluation
von: Shen, Chengyu, et al.
Veröffentlicht: (2026)
von: Shen, Chengyu, et al.
Veröffentlicht: (2026)
FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries
von: You, Qijie, et al.
Veröffentlicht: (2026)
von: You, Qijie, et al.
Veröffentlicht: (2026)
When grandparents step back: Fertility intentions and policy responses amid delayed retirement
von: Chen, He, et al.
Veröffentlicht: (2025)
von: Chen, He, et al.
Veröffentlicht: (2025)
Gradual Learning: Optimizing Fine-Tuning with Partially Mastered Knowledge in Large Language Models
von: Li, Bozhou, et al.
Veröffentlicht: (2024)
von: Li, Bozhou, et al.
Veröffentlicht: (2024)
Strong and weak well-posedness of McKean-Vlasov SDEs driven by $α$-stable processes under unified condition
von: Hao, Zimo
Veröffentlicht: (2025)
von: Hao, Zimo
Veröffentlicht: (2025)
Topic Over Source: The Key to Effective Data Mixing for Language Models Pre-training
von: Peng, Jiahui, et al.
Veröffentlicht: (2025)
von: Peng, Jiahui, et al.
Veröffentlicht: (2025)
FlipVQA: Scaling Multi-modal Instruction Tuning via Textbook-to-Knowledge Synthesis
von: Wong, Zhen Hao, et al.
Veröffentlicht: (2025)
von: Wong, Zhen Hao, et al.
Veröffentlicht: (2025)
The Roles of T Cells in the Development of Metabolic Dysfunction‐Associated Steatohepatitis
von: Zhifa Ge, et al.
Veröffentlicht: (2025)
von: Zhifa Ge, et al.
Veröffentlicht: (2025)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
von: Liu, Zhou, et al.
Veröffentlicht: (2025)
von: Liu, Zhou, et al.
Veröffentlicht: (2025)
MathClean: A Benchmark for Synthetic Mathematical Data Cleaning
von: Liang, Hao, et al.
Veröffentlicht: (2025)
von: Liang, Hao, et al.
Veröffentlicht: (2025)
Changes of filial responsibility norms under public long‐term care insurance in China
von: Qifeng Ma, et al.
Veröffentlicht: (2025)
von: Qifeng Ma, et al.
Veröffentlicht: (2025)
FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech
von: Ma, Linhan, et al.
Veröffentlicht: (2025)
von: Ma, Linhan, et al.
Veröffentlicht: (2025)
Deuteron gravitational form factors: exchange currents
von: He, Fangcheng, et al.
Veröffentlicht: (2024)
von: He, Fangcheng, et al.
Veröffentlicht: (2024)
Threshold photo-production of $J/Ψ$ off light nuclei
von: He, Fangcheng, et al.
Veröffentlicht: (2024)
von: He, Fangcheng, et al.
Veröffentlicht: (2024)
Helium-4 gravitational form factors: exchange currents
von: He, Fangcheng, et al.
Veröffentlicht: (2024)
von: He, Fangcheng, et al.
Veröffentlicht: (2024)
Gravitational form factors of light nuclei: Impulse approximation
von: He, Fangcheng, et al.
Veröffentlicht: (2023)
von: He, Fangcheng, et al.
Veröffentlicht: (2023)
In Situ Construction of Hollow Coral‐Like Porous S‐Doped g‐C3N4/ZnIn2S4 S‐Scheme Heterojunction for Efficient Photocatalytic Hydrogen Evolution
von: Tianyu Wang, et al.
Veröffentlicht: (2024)
von: Tianyu Wang, et al.
Veröffentlicht: (2024)
LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning
von: Wong, Zhen Hao, et al.
Veröffentlicht: (2025)
von: Wong, Zhen Hao, et al.
Veröffentlicht: (2025)
Improving Time Series Classification with Representation Soft Label Smoothing
von: Ma, Hengyi, et al.
Veröffentlicht: (2024)
von: Ma, Hengyi, et al.
Veröffentlicht: (2024)
SDEs with supercritical distributional drifts
von: Hao, Zimo, et al.
Veröffentlicht: (2023)
von: Hao, Zimo, et al.
Veröffentlicht: (2023)
SDE driven by cylindrical $α$-stable process with distributional drift
von: Hao, Zimo, et al.
Veröffentlicht: (2023)
von: Hao, Zimo, et al.
Veröffentlicht: (2023)
Convergence rate of the Euler-Maruyama scheme to density dependent SDEs driven by $α$-stable additive noise
von: Song, Ke, et al.
Veröffentlicht: (2024)
von: Song, Ke, et al.
Veröffentlicht: (2024)
Euler--Maruyama scheme for $α$-stable SDE with distributional drift
von: Hao, Zimo, et al.
Veröffentlicht: (2026)
von: Hao, Zimo, et al.
Veröffentlicht: (2026)
An optimized perfusate for enhanced rat ex vivo lung perfusion and lung transplant models
von: Jie Zhang, et al.
Veröffentlicht: (2025)
von: Jie Zhang, et al.
Veröffentlicht: (2025)
Generative Giants, Retrieval Weaklings: Why do Multimodal Large Language Models Fail at Multimodal Retrieval?
von: Feng, Hengyi, et al.
Veröffentlicht: (2025)
von: Feng, Hengyi, et al.
Veröffentlicht: (2025)
Synth-Empathy: Towards High-Quality Synthetic Empathy Data
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
CLIP-based Synergistic Knowledge Transfer for Text-based Person Retrieval
von: Liu, Yating, et al.
Veröffentlicht: (2023)
von: Liu, Yating, et al.
Veröffentlicht: (2023)
Formal Logic Enabled Personalized Federated Learning Through Property Inference
von: An, Ziyan, et al.
Veröffentlicht: (2024)
von: An, Ziyan, et al.
Veröffentlicht: (2024)
Towards Reliable Neural Optimizers: Permutation-Equivariant Neural Approximation in Dynamic Data Driven Applications Systems
von: Li, Meiyi, et al.
Veröffentlicht: (2025)
von: Li, Meiyi, et al.
Veröffentlicht: (2025)
Boosting Few-Shot Segmentation via Instance-Aware Data Augmentation and Local Consensus Guided Cross Attention
von: Guo, Li, et al.
Veröffentlicht: (2024)
von: Guo, Li, et al.
Veröffentlicht: (2024)
FlexQ: Efficient Post-training INT6 Quantization for LLM Serving via Algorithm-System Co-Design
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
von: Liang, Hao, et al.
Veröffentlicht: (2025) -
Towards Next-Generation LLM Training: From the Data-Centric Perspective
von: Liang, Hao, et al.
Veröffentlicht: (2026) -
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos
von: Feng, Hengyi, et al.
Veröffentlicht: (2026) -
Let's Verify Math Questions Step by Step
von: Shen, Chengyu, et al.
Veröffentlicht: (2025) -
Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
von: Ma, Lu, et al.
Veröffentlicht: (2025)