GraphGen: Enhancing Supervised Fine-Tuning for LLMs with Knowledge-Driven Synthetic Data Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Zihong, Jiang, Wanli, Li, Jinzhe, Yuan, Zhonghang, Kong, Huanjun, Ouyang, Wanli, Dong, Nanqing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains
by: Yuan, Zhonghang, et al.
Published: (2026)
by: Yuan, Zhonghang, et al.
Published: (2026)
ROGRAG: A Robustly Optimized GraphRAG Framework
by: Wang, Zhefan, et al.
Published: (2025)
by: Wang, Zhefan, et al.
Published: (2025)
SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science
by: Ying, Jie, et al.
Published: (2025)
by: Ying, Jie, et al.
Published: (2025)
Labeling supervised fine-tuning data with the scaling law
by: Kong, Huanjun
Published: (2024)
by: Kong, Huanjun
Published: (2024)
GraphGen+: Advancing Distributed Subgraph Generation and Graph Learning On Industrial Graphs
by: Jin, Yue, et al.
Published: (2025)
by: Jin, Yue, et al.
Published: (2025)
AI-Driven Automation Can Become the Foundation of Next-Era Science of Science Research
by: Chen, Renqi, et al.
Published: (2025)
by: Chen, Renqi, et al.
Published: (2025)
Towards Efficient and Intelligent Laser Weeding: Method and Dataset for Weed Stem Detection
by: Liu, Dingning, et al.
Published: (2025)
by: Liu, Dingning, et al.
Published: (2025)
Retrieval is Not Enough: Enhancing RAG Reasoning through Test-Time Critique and Optimization
by: Wei, Jiaqi, et al.
Published: (2025)
by: Wei, Jiaqi, et al.
Published: (2025)
Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised Fine-Tuning
by: Singh, Navan Preet, et al.
Published: (2026)
by: Singh, Navan Preet, et al.
Published: (2026)
Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System
by: Su, Haoyang, et al.
Published: (2024)
by: Su, Haoyang, et al.
Published: (2024)
Fine-Tuning LLMs for Report Summarization: Analysis on Supervised and Unsupervised Data
by: Rallapalli, Swati, et al.
Published: (2025)
by: Rallapalli, Swati, et al.
Published: (2025)
Rethinking Data Selection for Supervised Fine-Tuning
by: Shen, Ming
Published: (2024)
by: Shen, Ming
Published: (2024)
Generating Logically Consistent Synthetic Supply Chain Data with LLM-Driven Knowledge Graph Reasoning
by: Long, Yunbo, et al.
Published: (2026)
by: Long, Yunbo, et al.
Published: (2026)
Towards Pedagogical LLMs with Supervised Fine Tuning for Computing Education
by: Vassar, Alexandra, et al.
Published: (2024)
by: Vassar, Alexandra, et al.
Published: (2024)
Mind the Gap: Data Rewriting for Stable Off-Policy Supervised Fine-Tuning
by: Zhao, Shiwan, et al.
Published: (2025)
by: Zhao, Shiwan, et al.
Published: (2025)
Beyond Reasoning: Reinforcement Learning Unlocks Parametric Knowledge in LLMs
by: Yang, Wanli, et al.
Published: (2026)
by: Yang, Wanli, et al.
Published: (2026)
Semantic Loss Guided Data Efficient Supervised Fine Tuning for Safe Responses in LLMs
by: Lu, Yuxiao, et al.
Published: (2024)
by: Lu, Yuxiao, et al.
Published: (2024)
KG-FIT: Knowledge Graph Fine-Tuning Upon Open-World Knowledge
by: Jiang, Pengcheng, et al.
Published: (2024)
by: Jiang, Pengcheng, et al.
Published: (2024)
MOSLIM:Align with diverse preferences in prompts through reward classification
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Repurposing Synthetic Data for Fine-grained Search Agent Supervision
by: Zhao, Yida, et al.
Published: (2025)
by: Zhao, Yida, et al.
Published: (2025)
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey
by: Long, Lin, et al.
Published: (2024)
by: Long, Lin, et al.
Published: (2024)
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
by: Gekhman, Zorik, et al.
Published: (2024)
by: Gekhman, Zorik, et al.
Published: (2024)
Supervised In-Context Fine-Tuning for Generative Sequence Labeling
by: Dukić, David, et al.
Published: (2025)
by: Dukić, David, et al.
Published: (2025)
Synthetic Eggs in Many Baskets: The Impact of Synthetic Data Diversity on LLM Fine-Tuning
by: Schaffelder, Max, et al.
Published: (2025)
by: Schaffelder, Max, et al.
Published: (2025)
Augmented Fine-Tuned LLMs for Enhanced Recruitment Automation
by: Younes, Mohamed T., et al.
Published: (2025)
by: Younes, Mohamed T., et al.
Published: (2025)
One-Token Rollout: Guiding Supervised Fine-Tuning of LLMs with Policy Gradient
by: Ming, Rui, et al.
Published: (2025)
by: Ming, Rui, et al.
Published: (2025)
Blinded by Generated Contexts: How Language Models Merge Generated and Retrieved Contexts When Knowledge Conflicts?
by: Tan, Hexiang, et al.
Published: (2024)
by: Tan, Hexiang, et al.
Published: (2024)
Supervised Fine-Tuning or In-Context Learning? Evaluating LLMs for Clinical NER
by: Baroian, Andrei
Published: (2025)
by: Baroian, Andrei
Published: (2025)
Supervised Fine-Tuning LLMs to Behave as Pedagogical Agents in Programming Education
by: Ross, Emily, et al.
Published: (2025)
by: Ross, Emily, et al.
Published: (2025)
SHAPE-IT: Exploring Text-to-Shape-Display for Generative Shape-Changing Behaviors with LLMs
by: Qian, Wanli, et al.
Published: (2024)
by: Qian, Wanli, et al.
Published: (2024)
Threshold Filtering Packing for Supervised Fine-Tuning: Training Related Samples within Packs
by: Dong, Jiancheng, et al.
Published: (2024)
by: Dong, Jiancheng, et al.
Published: (2024)
Injecting New Knowledge into Large Language Models via Supervised Fine-Tuning
by: Mecklenburg, Nick, et al.
Published: (2024)
by: Mecklenburg, Nick, et al.
Published: (2024)
Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs
by: Ovadia, Oded, et al.
Published: (2023)
by: Ovadia, Oded, et al.
Published: (2023)
ASVRI-Legal: Fine-Tuning LLMs with Retrieval Augmented Generation for Enhanced Legal Regulation
by: Octadion, One, et al.
Published: (2025)
by: Octadion, One, et al.
Published: (2025)
Enhancing Agentic Textual Graph Retrieval with Synthetic Stepwise Supervision
by: Chang, Ge, et al.
Published: (2025)
by: Chang, Ge, et al.
Published: (2025)
FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training
by: Yu, Hongzhou, et al.
Published: (2025)
by: Yu, Hongzhou, et al.
Published: (2025)
Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning
by: Kopiczko, Dawid J., et al.
Published: (2026)
by: Kopiczko, Dawid J., et al.
Published: (2026)
Token Cleaning: Fine-Grained Data Selection for LLM Supervised Fine-Tuning
by: Pang, Jinlong, et al.
Published: (2025)
by: Pang, Jinlong, et al.
Published: (2025)
Anchored Supervised Fine-Tuning
by: Zhu, He, et al.
Published: (2025)
by: Zhu, He, et al.
Published: (2025)
Toward Understanding BERT-Like Pre-Training for DNA Foundation Models
by: Liang, Chaoqi, et al.
Published: (2023)
by: Liang, Chaoqi, et al.
Published: (2023)
Similar Items
-
Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains
by: Yuan, Zhonghang, et al.
Published: (2026) -
ROGRAG: A Robustly Optimized GraphRAG Framework
by: Wang, Zhefan, et al.
Published: (2025) -
SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science
by: Ying, Jie, et al.
Published: (2025) -
Labeling supervised fine-tuning data with the scaling law
by: Kong, Huanjun
Published: (2024) -
GraphGen+: Advancing Distributed Subgraph Generation and Graph Learning On Industrial Graphs
by: Jin, Yue, et al.
Published: (2025)