DataArc-SynData-Toolkit: A Unified Closed-Loop Framework for Multi-Path, Multimodal, and Multilingual Data Synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Zhichao, Yang, Cehao, Zhou, Hao, Wu, Xiaojun, Li, Huajie, Jiang, Xuhui, Xu, Chengjin, Wang, Yuanzhuo, Guo, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Synthesize-on-Graph: Knowledgeable Synthetic Data Generation for Continue Pre-training of Large Language Models
von: Ma, Shengjie, et al.
Veröffentlicht: (2025)
von: Ma, Shengjie, et al.
Veröffentlicht: (2025)
Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
Continual Pretraining on Encrypted Synthetic Data for Privacy-Preserving LLMs
von: Liu, Honghao, et al.
Veröffentlicht: (2026)
von: Liu, Honghao, et al.
Veröffentlicht: (2026)
LongFaith: Enhancing Long-Context Reasoning in LLMs with Faithful Synthetic Data
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
A Survey on Large Language Model Hallucination via a Creativity Perspective
von: Jiang, Xuhui, et al.
Veröffentlicht: (2024)
von: Jiang, Xuhui, et al.
Veröffentlicht: (2024)
Unlocking the Power of Large Language Models for Entity Alignment
von: Jiang, Xuhui, et al.
Veröffentlicht: (2024)
von: Jiang, Xuhui, et al.
Veröffentlicht: (2024)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
von: Shi, Zhichao, et al.
Veröffentlicht: (2025)
von: Shi, Zhichao, et al.
Veröffentlicht: (2025)
Context Graph
von: Xu, Chengjin, et al.
Veröffentlicht: (2024)
von: Xu, Chengjin, et al.
Veröffentlicht: (2024)
Think-on-Graph 3.0: Efficient and Adaptive LLM Reasoning on Heterogeneous Graphs via Multi-Agent Dual-Evolving Context Retrieval
von: Wu, Xiaojun, et al.
Veröffentlicht: (2025)
von: Wu, Xiaojun, et al.
Veröffentlicht: (2025)
GraphSearch: An Agentic Deep Searching Workflow for Graph Retrieval-Augmented Generation
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
Retrieval, Reasoning, Re-ranking: A Context-Enriched Framework for Knowledge Graph Completion
von: Li, Muzhi, et al.
Veröffentlicht: (2024)
von: Li, Muzhi, et al.
Veröffentlicht: (2024)
Toward Practical Entity Alignment Method Design: Insights from New Highly Heterogeneous Knowledge Graph Datasets
von: Jiang, Xuhui, et al.
Veröffentlicht: (2023)
von: Jiang, Xuhui, et al.
Veröffentlicht: (2023)
Financial Knowledge Large Language Model
von: Yang, Cehao, et al.
Veröffentlicht: (2024)
von: Yang, Cehao, et al.
Veröffentlicht: (2024)
Think-on-Graph 2.0: Deep and Faithful Large Language Model Reasoning with Knowledge-guided Retrieval Augmented Generation
von: Ma, Shengjie, et al.
Veröffentlicht: (2024)
von: Ma, Shengjie, et al.
Veröffentlicht: (2024)
Conflicts Make Large Reasoning Models Vulnerable to Attacks
von: Liu, Honghao, et al.
Veröffentlicht: (2026)
von: Liu, Honghao, et al.
Veröffentlicht: (2026)
Context-aware Inductive Knowledge Graph Completion with Latent Type Constraints and Subgraph Reasoning
von: Li, Muzhi, et al.
Veröffentlicht: (2024)
von: Li, Muzhi, et al.
Veröffentlicht: (2024)
On the Evolution of Knowledge Graphs: A Survey and Perspective
von: Jiang, Xuhui, et al.
Veröffentlicht: (2023)
von: Jiang, Xuhui, et al.
Veröffentlicht: (2023)
An LLM-Assisted Toolkit for Inspectable Multimodal Emotion Data Annotation
von: Kuang, Zheyuan, et al.
Veröffentlicht: (2026)
von: Kuang, Zheyuan, et al.
Veröffentlicht: (2026)
Financial Wind Tunnel: A Retrieval-Augmented Market Simulator
von: Cao, Bokai, et al.
Veröffentlicht: (2025)
von: Cao, Bokai, et al.
Veröffentlicht: (2025)
Data for Case Study, State Space Boundary, and Comparison of TV-DSSR model and TI-DSSR model
von: Li, Chengjin
Veröffentlicht: (2025)
von: Li, Chengjin
Veröffentlicht: (2025)
EvoSyn: Generalizable Evolutionary Data Synthesis for Verifiable Learning
von: Du, He, et al.
Veröffentlicht: (2025)
von: Du, He, et al.
Veröffentlicht: (2025)
Syn-GRPO: Self-Evolving Data Synthesis for MLLM Perception Reasoning
von: Huang, Qihan, et al.
Veröffentlicht: (2025)
von: Huang, Qihan, et al.
Veröffentlicht: (2025)
Empowering Bridge Digital Twins by Bridging the Data Gap with a Unified Synthesis Framework
von: Wang, Wang, et al.
Veröffentlicht: (2025)
von: Wang, Wang, et al.
Veröffentlicht: (2025)
RV-Syn: Rational and Verifiable Mathematical Reasoning Data Synthesis based on Structured Function Library
von: Wang, Jiapeng, et al.
Veröffentlicht: (2025)
von: Wang, Jiapeng, et al.
Veröffentlicht: (2025)
SynSHRP2: A Synthetic Multimodal Benchmark for Driving Safety-critical Events Derived from Real-world Driving Data
von: Shi, Liang, et al.
Veröffentlicht: (2025)
von: Shi, Liang, et al.
Veröffentlicht: (2025)
ReTabSyn: Realistic Tabular Data Synthesis via Reinforcement Learning
von: Lin, Xiaofeng, et al.
Veröffentlicht: (2026)
von: Lin, Xiaofeng, et al.
Veröffentlicht: (2026)
K-Syn: K-space Data Synthesis in Ultra Low-data Regimes
von: Yu, Guan, et al.
Veröffentlicht: (2025)
von: Yu, Guan, et al.
Veröffentlicht: (2025)
Closing the Data Loop: Using OpenDataArena to Engineer Superior Training Datasets
von: Gao, Xin, et al.
Veröffentlicht: (2025)
von: Gao, Xin, et al.
Veröffentlicht: (2025)
Data-Driven Predictive Control Using Closed-Loop Data: An Instrumental Variable Approach
von: Wang, Yibo, et al.
Veröffentlicht: (2023)
von: Wang, Yibo, et al.
Veröffentlicht: (2023)
LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
von: Zhang, Kangning, et al.
Veröffentlicht: (2025)
von: Zhang, Kangning, et al.
Veröffentlicht: (2025)
RETuning: Upgrading Inference-Time Scaling for Stock Movement Prediction with Large Language Models
von: Lin, Xueyuan, et al.
Veröffentlicht: (2025)
von: Lin, Xueyuan, et al.
Veröffentlicht: (2025)
SynSym: A Synthetic Data Generation Framework for Psychiatric Symptom Identification
von: Kang, Migyeong, et al.
Veröffentlicht: (2026)
von: Kang, Migyeong, et al.
Veröffentlicht: (2026)
HeteroFedSyn: Differentially Private Tabular Data Synthesis for Heterogeneous Federated Settings
von: Li, Xiaochen, et al.
Veröffentlicht: (2026)
von: Li, Xiaochen, et al.
Veröffentlicht: (2026)
A Survey on LLM-as-a-Judge
von: Gu, Jiawei, et al.
Veröffentlicht: (2024)
von: Gu, Jiawei, et al.
Veröffentlicht: (2024)
SQL-R1: Training Natural Language to SQL Reasoning Model By Reinforcement Learning
von: Ma, Peixian, et al.
Veröffentlicht: (2025)
von: Ma, Peixian, et al.
Veröffentlicht: (2025)
ProofWala: A Framework for Multilingual Proof Data Synthesis and Theorem-Proving
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025)
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025)
FICO: Finite-Horizon Closed-Loop Factorization for Unified Multi-Agent Path Finding
von: Li, Jiarui, et al.
Veröffentlicht: (2025)
von: Li, Jiarui, et al.
Veröffentlicht: (2025)
Functional Consistency of LLM Code Embeddings: A Self-Evolving Data Synthesis Framework for Benchmarking
von: Li, Zhuohao, et al.
Veröffentlicht: (2025)
von: Li, Zhuohao, et al.
Veröffentlicht: (2025)
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
SynQP: A Framework and Metrics for Evaluating the Quality and Privacy Risk of Synthetic Data
von: Hu, Bing, et al.
Veröffentlicht: (2026)
von: Hu, Bing, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Synthesize-on-Graph: Knowledgeable Synthetic Data Generation for Continue Pre-training of Large Language Models
von: Ma, Shengjie, et al.
Veröffentlicht: (2025) -
Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning
von: Yang, Cehao, et al.
Veröffentlicht: (2025) -
Continual Pretraining on Encrypted Synthetic Data for Privacy-Preserving LLMs
von: Liu, Honghao, et al.
Veröffentlicht: (2026) -
LongFaith: Enhancing Long-Context Reasoning in LLMs with Faithful Synthetic Data
von: Yang, Cehao, et al.
Veröffentlicht: (2025) -
A Survey on Large Language Model Hallucination via a Creativity Perspective
von: Jiang, Xuhui, et al.
Veröffentlicht: (2024)