Unleashing the Power of Data Tsunami: A Comprehensive Survey on Data Assessment and Selection for Instruction Tuning of Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qin, Yulei, Yang, Yuncheng, Guo, Pengcheng, Li, Gang, Shao, Hang, Shi, Yuchen, Xu, Zihan, Gu, Yun, Li, Ke, Sun, Xing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Open Knowledge for Advancing Task Expertise in Large Language Models
von: Yang, Yuncheng, et al.
Veröffentlicht: (2024)
von: Yang, Yuncheng, et al.
Veröffentlicht: (2024)
Incentivizing Reasoning for Advanced Instruction-Following of Large Language Models
von: Qin, Yulei, et al.
Veröffentlicht: (2025)
von: Qin, Yulei, et al.
Veröffentlicht: (2025)
SCAR: Data Selection via Style Consistency-Aware Response Ranking for Efficient Instruction-Tuning of Large Language Models
von: Li, Zhuang, et al.
Veröffentlicht: (2024)
von: Li, Zhuang, et al.
Veröffentlicht: (2024)
A Survey on Data Selection for LLM Instruction Tuning
von: Zhang, Bolin, et al.
Veröffentlicht: (2024)
von: Zhang, Bolin, et al.
Veröffentlicht: (2024)
RoRecomp: Enhancing Reasoning Efficiency via Rollout Response Recomposition in Reinforcement Learning
von: Li, Gang, et al.
Veröffentlicht: (2025)
von: Li, Gang, et al.
Veröffentlicht: (2025)
RESTORE: Towards Feature Shift for Vision-Language Prompt Learning
von: Yang, Yuncheng, et al.
Veröffentlicht: (2024)
von: Yang, Yuncheng, et al.
Veröffentlicht: (2024)
Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
LTD-Bench: Evaluating Large Language Models by Letting Them Draw
von: Lin, Liuhao, et al.
Veröffentlicht: (2025)
von: Lin, Liuhao, et al.
Veröffentlicht: (2025)
Data Selection for Multi-turn Dialogue Instruction Tuning
von: Li, Bo, et al.
Veröffentlicht: (2026)
von: Li, Bo, et al.
Veröffentlicht: (2026)
Less is More: High-value Data Selection for Visual Instruction Tuning
von: Liu, Zikang, et al.
Veröffentlicht: (2024)
von: Liu, Zikang, et al.
Veröffentlicht: (2024)
Large-Scale Data Selection for Instruction Tuning
von: Ivison, Hamish, et al.
Veröffentlicht: (2025)
von: Ivison, Hamish, et al.
Veröffentlicht: (2025)
Beyond Similarity: A Gradient-based Graph Method for Instruction Tuning Data Selection
von: Zhao, Yang, et al.
Veröffentlicht: (2025)
von: Zhao, Yang, et al.
Veröffentlicht: (2025)
Importance-Aware Data Selection for Efficient LLM Instruction Tuning
von: Jiang, Tingyu, et al.
Veröffentlicht: (2025)
von: Jiang, Tingyu, et al.
Veröffentlicht: (2025)
From Tags to Trees: Structuring Fine-Grained Knowledge for Controllable Data Selection in LLM Instruction Tuning
von: Niu, Zihan, et al.
Veröffentlicht: (2026)
von: Niu, Zihan, et al.
Veröffentlicht: (2026)
Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data
von: Li, Jiajie, et al.
Veröffentlicht: (2025)
von: Li, Jiajie, et al.
Veröffentlicht: (2025)
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
von: Liu, Wei, et al.
Veröffentlicht: (2023)
von: Liu, Wei, et al.
Veröffentlicht: (2023)
IterSelectTune: An Iterative Training Framework for Efficient Instruction-Tuning Data Selection
von: Song, Jielin, et al.
Veröffentlicht: (2024)
von: Song, Jielin, et al.
Veröffentlicht: (2024)
CoIDO: Efficient Data Selection for Visual Instruction Tuning via Coupled Importance-Diversity Optimization
von: Yan, Yichen, et al.
Veröffentlicht: (2025)
von: Yan, Yichen, et al.
Veröffentlicht: (2025)
Training-Free Group Relative Policy Optimization
von: Cai, Yuzheng, et al.
Veröffentlicht: (2025)
von: Cai, Yuzheng, et al.
Veröffentlicht: (2025)
Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction Tuning
von: Safaei, Bardia, et al.
Veröffentlicht: (2025)
von: Safaei, Bardia, et al.
Veröffentlicht: (2025)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
PACIT: Unlocking the Power of Examples for Better In-Context Instruction Tuning
von: Xue, Tianci, et al.
Veröffentlicht: (2023)
von: Xue, Tianci, et al.
Veröffentlicht: (2023)
Star-Agents: Automatic Data Optimization with LLM Agents for Instruction Tuning
von: Zhou, Hang, et al.
Veröffentlicht: (2024)
von: Zhou, Hang, et al.
Veröffentlicht: (2024)
GIST: Targeted Data Selection for Instruction Tuning via Coupled Optimization Geometry
von: Min, Guanghui, et al.
Veröffentlicht: (2026)
von: Min, Guanghui, et al.
Veröffentlicht: (2026)
CUARewardBench: A Benchmark for Evaluating Reward Models on Computer-using Agent
von: Lin, Haojia, et al.
Veröffentlicht: (2025)
von: Lin, Haojia, et al.
Veröffentlicht: (2025)
Unleashing the Power of Self-Supervised Image Denoising: A Comprehensive Review
von: Zhang, Dan, et al.
Veröffentlicht: (2023)
von: Zhang, Dan, et al.
Veröffentlicht: (2023)
SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents
von: Cai, Shaofei, et al.
Veröffentlicht: (2025)
von: Cai, Shaofei, et al.
Veröffentlicht: (2025)
TAGCOS: Task-agnostic Gradient Clustered Coreset Selection for Instruction Tuning Data
von: Zhang, Jipeng, et al.
Veröffentlicht: (2024)
von: Zhang, Jipeng, et al.
Veröffentlicht: (2024)
From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning
von: Li, Ming, et al.
Veröffentlicht: (2023)
von: Li, Ming, et al.
Veröffentlicht: (2023)
Instruction Mining: Instruction Data Selection for Tuning Large Language Models
von: Cao, Yihan, et al.
Veröffentlicht: (2023)
von: Cao, Yihan, et al.
Veröffentlicht: (2023)
LangGPS: Language Separability Guided Data Pre-Selection for Joint Multilingual Instruction Tuning
von: Ye, Yangfan, et al.
Veröffentlicht: (2025)
von: Ye, Yangfan, et al.
Veröffentlicht: (2025)
LESS: Selecting Influential Data for Targeted Instruction Tuning
von: Xia, Mengzhou, et al.
Veröffentlicht: (2024)
von: Xia, Mengzhou, et al.
Veröffentlicht: (2024)
Neuron-Aware Data Selection In Instruction Tuning For Large Language Models
von: Chen, Xin, et al.
Veröffentlicht: (2026)
von: Chen, Xin, et al.
Veröffentlicht: (2026)
RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection
von: Yang, Yixin, et al.
Veröffentlicht: (2025)
von: Yang, Yixin, et al.
Veröffentlicht: (2025)
Data Diversity Matters for Robust Instruction Tuning
von: Bukharin, Alexander, et al.
Veröffentlicht: (2023)
von: Bukharin, Alexander, et al.
Veröffentlicht: (2023)
MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space
von: Chen, Yicheng, et al.
Veröffentlicht: (2025)
von: Chen, Yicheng, et al.
Veröffentlicht: (2025)
SC-Tune: Unleashing Self-Consistent Referential Comprehension in Large Vision Language Models
von: Yue, Tongtian, et al.
Veröffentlicht: (2024)
von: Yue, Tongtian, et al.
Veröffentlicht: (2024)
Instruction Data Selection via Answer Divergence
von: Li, Bo, et al.
Veröffentlicht: (2026)
von: Li, Bo, et al.
Veröffentlicht: (2026)
Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Leveraging Open Knowledge for Advancing Task Expertise in Large Language Models
von: Yang, Yuncheng, et al.
Veröffentlicht: (2024) -
Incentivizing Reasoning for Advanced Instruction-Following of Large Language Models
von: Qin, Yulei, et al.
Veröffentlicht: (2025) -
SCAR: Data Selection via Style Consistency-Aware Response Ranking for Efficient Instruction-Tuning of Large Language Models
von: Li, Zhuang, et al.
Veröffentlicht: (2024) -
A Survey on Data Selection for LLM Instruction Tuning
von: Zhang, Bolin, et al.
Veröffentlicht: (2024) -
RoRecomp: Enhancing Reasoning Efficiency via Rollout Response Recomposition in Reinforcement Learning
von: Li, Gang, et al.
Veröffentlicht: (2025)