What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Yuchang, Zhong, Huazhen, Lin, Qunshu, Wei, Haotong, Sun, Xiaolong, Yu, Zixuan, Liu, Minghao, Zheng, Zibin, Chen, Liang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GT-SNT: A Linear-Time Transformer for Large-Scale Graphs via Spiking Node Tokenization
by: Zhang, Huizhe, et al.
Published: (2025)
by: Zhang, Huizhe, et al.
Published: (2025)
Fair Graph Representation Learning via Sensitive Attribute Disentanglement
by: Zhu, Yuchang, et al.
Published: (2024)
by: Zhu, Yuchang, et al.
Published: (2024)
Measuring Diversity in Synthetic Datasets
by: Zhu, Yuchang, et al.
Published: (2025)
by: Zhu, Yuchang, et al.
Published: (2025)
One Fits All: Learning Fair Graph Neural Networks for Various Sensitive Attributes
by: Zhu, Yuchang, et al.
Published: (2024)
by: Zhu, Yuchang, et al.
Published: (2024)
SaGIF: Improving Individual Fairness in Graph Neural Networks via Similarity Encoding
by: Zhu, Yuchang, et al.
Published: (2025)
by: Zhu, Yuchang, et al.
Published: (2025)
Oversmoothing: A Nightmare for Graph Contrastive Learning?
by: Li, Jintang, et al.
Published: (2023)
by: Li, Jintang, et al.
Published: (2023)
Are Large Language Models In-Context Graph Learners?
by: Li, Jintang, et al.
Published: (2025)
by: Li, Jintang, et al.
Published: (2025)
Exploring Selective Layer Fine-Tuning in Federated Learning
by: Sun, Yuchang, et al.
Published: (2024)
by: Sun, Yuchang, et al.
Published: (2024)
Low-rank Attention Side-Tuning for Parameter-Efficient Fine-Tuning
by: Tang, Ningyuan, et al.
Published: (2024)
by: Tang, Ningyuan, et al.
Published: (2024)
Token Cleaning: Fine-Grained Data Selection for LLM Supervised Fine-Tuning
by: Pang, Jinlong, et al.
Published: (2025)
by: Pang, Jinlong, et al.
Published: (2025)
On the Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models
by: Wang, Shumin, et al.
Published: (2026)
by: Wang, Shumin, et al.
Published: (2026)
Generative adversarial learning with optimal input dimension and its adaptive generator architecture
by: Tan, Zhiyao, et al.
Published: (2024)
by: Tan, Zhiyao, et al.
Published: (2024)
Data Diversity Matters for Robust Instruction Tuning
by: Bukharin, Alexander, et al.
Published: (2023)
by: Bukharin, Alexander, et al.
Published: (2023)
The Long-Term Effects of Data Selection in LLM Fine-Tuning
by: Yang, Yuxin, et al.
Published: (2026)
by: Yang, Yuxin, et al.
Published: (2026)
Synthetic Eggs in Many Baskets: The Impact of Synthetic Data Diversity on LLM Fine-Tuning
by: Schaffelder, Max, et al.
Published: (2025)
by: Schaffelder, Max, et al.
Published: (2025)
MindGYM: What Matters in Question Synthesis for Thinking-Centric Fine-Tuning?
by: Xu, Zhe, et al.
Published: (2025)
by: Xu, Zhe, et al.
Published: (2025)
What Matters in Data for DPO?
by: Pan, Yu, et al.
Published: (2025)
by: Pan, Yu, et al.
Published: (2025)
Adaptive Defense against Harmful Fine-Tuning for Large Language Models via Bayesian Data Scheduler
by: Hu, Zixuan, et al.
Published: (2025)
by: Hu, Zixuan, et al.
Published: (2025)
Learn How to Query from Unlabeled Data Streams in Federated Learning
by: Sun, Yuchang, et al.
Published: (2024)
by: Sun, Yuchang, et al.
Published: (2024)
Mixture-of-Skills: Learning to Optimize Data Usage for Fine-Tuning Large Language Models
by: Wu, Minghao, et al.
Published: (2024)
by: Wu, Minghao, et al.
Published: (2024)
A limiter-based approach to construct high-order fully-discrete entropy stable explicit DG schemes for hyperbolic conservation laws
by: Liu, Yuchang, et al.
Published: (2026)
by: Liu, Yuchang, et al.
Published: (2026)
DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning
by: Zhang, Junbo, et al.
Published: (2026)
by: Zhang, Junbo, et al.
Published: (2026)
Olica: Efficient Structured Pruning of Large Language Models without Retraining
by: He, Jiujun, et al.
Published: (2025)
by: He, Jiujun, et al.
Published: (2025)
Augmenting Smart Contract Decompiler Output through Fine-grained Dependency Analysis and LLM-facilitated Semantic Recovery
by: Liao, Zeqin, et al.
Published: (2025)
by: Liao, Zeqin, et al.
Published: (2025)
AgentRaft: Automated Detection of Data Over-Exposure in LLM Agents
by: Lin, Yixi, et al.
Published: (2026)
by: Lin, Yixi, et al.
Published: (2026)
Why Tropical Cyclones Over Oceanic Cyclonic Eddies Can Be Intensified in Global Basins
by: Lingwei Wu, et al.
Published: (2026)
by: Lingwei Wu, et al.
Published: (2026)
MCTS-Refined CoT: High-Quality Fine-Tuning Data for LLM-Based Repository Issue Resolution
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
When Fine-Tuning Fails and when it Generalises: Role of Data Diversity and Mixed Training in LLM-based TTS
by: Purwar, Anupam, et al.
Published: (2026)
by: Purwar, Anupam, et al.
Published: (2026)
Secure LLM Fine-Tuning via Safety-Aware Probing
by: Wu, Chengcan, et al.
Published: (2025)
by: Wu, Chengcan, et al.
Published: (2025)
Improving time series estimation and prediction via transfer learning
by: Lin, Yuchang, et al.
Published: (2025)
by: Lin, Yuchang, et al.
Published: (2025)
A robust and scalable estimation for high-dimensional volatility models
by: Chen, Kejun, et al.
Published: (2025)
by: Chen, Kejun, et al.
Published: (2025)
Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning
by: Zheng, Han, et al.
Published: (2026)
by: Zheng, Han, et al.
Published: (2026)
Scaling Data Diversity for Fine-Tuning Language Models in Human Alignment
by: Song, Feifan, et al.
Published: (2024)
by: Song, Feifan, et al.
Published: (2024)
On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
by: Zhang, Wenhao, et al.
Published: (2025)
by: Zhang, Wenhao, et al.
Published: (2025)
Easy Dataset: A Unified and Extensible Framework for Synthesizing LLM Fine-Tuning Data from Unstructured Documents
by: Miao, Ziyang, et al.
Published: (2025)
by: Miao, Ziyang, et al.
Published: (2025)
Rank Also Matters: Hierarchical Configuration for Mixture of Adapter Experts in LLM Fine-Tuning
by: Cong, Peizhuang, et al.
Published: (2025)
by: Cong, Peizhuang, et al.
Published: (2025)
Objaverse++: Curated 3D Object Dataset with Quality Annotations
by: Lin, Chendi, et al.
Published: (2025)
by: Lin, Chendi, et al.
Published: (2025)
Diverse and Fine-Grained Instruction-Following Ability Exploration with Synthetic Data
by: Gu, Zihui, et al.
Published: (2024)
by: Gu, Zihui, et al.
Published: (2024)
Detecting Dataset Abuse in Fine-Tuning Stable Diffusion Models for Text-to-Image Synthesis
by: Wang, Songrui, et al.
Published: (2024)
by: Wang, Songrui, et al.
Published: (2024)
MoFO: Momentum-Filtered Optimizer for Mitigating Forgetting in LLM Fine-Tuning
by: Chen, Yupeng, et al.
Published: (2024)
by: Chen, Yupeng, et al.
Published: (2024)
Similar Items
-
GT-SNT: A Linear-Time Transformer for Large-Scale Graphs via Spiking Node Tokenization
by: Zhang, Huizhe, et al.
Published: (2025) -
Fair Graph Representation Learning via Sensitive Attribute Disentanglement
by: Zhu, Yuchang, et al.
Published: (2024) -
Measuring Diversity in Synthetic Datasets
by: Zhu, Yuchang, et al.
Published: (2025) -
One Fits All: Learning Fair Graph Neural Networks for Various Sensitive Attributes
by: Zhu, Yuchang, et al.
Published: (2024) -
SaGIF: Improving Individual Fairness in Graph Neural Networks via Similarity Encoding
by: Zhu, Yuchang, et al.
Published: (2025)