Importance-Aware Data Selection for Efficient LLM Instruction Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Tingyu, Li, Shen, Song, Yiyao, Zhang, Lan, Zhu, Hualei, Zhao, Yuan, Xu, Xiaohang, Taura, Kenjiro, Wang, Hao Henry |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Co-EPG: A Framework for Co-Evolution of Planning and Grounding in Autonomous GUI Agents
by: Zhao, Yuan, et al.
Published: (2025)
by: Zhao, Yuan, et al.
Published: (2025)
IterSelectTune: An Iterative Training Framework for Efficient Instruction-Tuning Data Selection
by: Song, Jielin, et al.
Published: (2024)
by: Song, Jielin, et al.
Published: (2024)
CoachLM: Automatic Instruction Revisions Improve the Data Quality in LLM Instruction Tuning
by: Liu, Yilun, et al.
Published: (2023)
by: Liu, Yilun, et al.
Published: (2023)
A Survey on Data Selection for LLM Instruction Tuning
by: Zhang, Bolin, et al.
Published: (2024)
by: Zhang, Bolin, et al.
Published: (2024)
Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning
by: Yang, Cehao, et al.
Published: (2025)
by: Yang, Cehao, et al.
Published: (2025)
Uncertainty-Aware Gradient Signal-to-Noise Data Selection for Instruction Tuning
by: Yuan, Zhihang, et al.
Published: (2026)
by: Yuan, Zhihang, et al.
Published: (2026)
Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
Neuron-Aware Data Selection In Instruction Tuning For Large Language Models
by: Chen, Xin, et al.
Published: (2026)
by: Chen, Xin, et al.
Published: (2026)
GIFT: Guided Importance-Aware Fine-Tuning for Diffusion Language Models
by: Xu, Guowei, et al.
Published: (2025)
by: Xu, Guowei, et al.
Published: (2025)
Unified Data Selection for LLM Reasoning
by: Li, Xiaoyuan, et al.
Published: (2026)
by: Li, Xiaoyuan, et al.
Published: (2026)
Steering at the Source: Style Modulation Heads for Robust Persona Control
by: Izawa, Yoshihiro, et al.
Published: (2026)
by: Izawa, Yoshihiro, et al.
Published: (2026)
GRAIT: Gradient-Driven Refusal-Aware Instruction Tuning for Effective Hallucination Mitigation
by: Zhu, Runchuan, et al.
Published: (2025)
by: Zhu, Runchuan, et al.
Published: (2025)
SCAR: Data Selection via Style Consistency-Aware Response Ranking for Efficient Instruction-Tuning of Large Language Models
by: Li, Zhuang, et al.
Published: (2024)
by: Li, Zhuang, et al.
Published: (2024)
From Tags to Trees: Structuring Fine-Grained Knowledge for Controllable Data Selection in LLM Instruction Tuning
by: Niu, Zihan, et al.
Published: (2026)
by: Niu, Zihan, et al.
Published: (2026)
Large-Scale Data Selection for Instruction Tuning
by: Ivison, Hamish, et al.
Published: (2025)
by: Ivison, Hamish, et al.
Published: (2025)
Llama-Mob: Instruction-Tuning Llama-3-8B Excels in City-Scale Mobility Prediction
by: Tang, Peizhi, et al.
Published: (2024)
by: Tang, Peizhi, et al.
Published: (2024)
RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection
by: Yang, Yixin, et al.
Published: (2025)
by: Yang, Yixin, et al.
Published: (2025)
GTaP: A GPU-Resident Fork-Join Task-Parallel Runtime with a Pragma-Based Interface
by: Maeda, Yuki, et al.
Published: (2026)
by: Maeda, Yuki, et al.
Published: (2026)
DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning
by: Zhang, Junbo, et al.
Published: (2026)
by: Zhang, Junbo, et al.
Published: (2026)
Skill-Aware Data Selection and Fine-Tuning for Data-Efficient Reasoning Distillation
by: Zhang, Lechen, et al.
Published: (2026)
by: Zhang, Lechen, et al.
Published: (2026)
IAPT: Instruction-Aware Prompt Tuning for Large Language Models
by: Zhu, Wei, et al.
Published: (2024)
by: Zhu, Wei, et al.
Published: (2024)
CrowdSelect: Synthetic Instruction Data Selection with Multi-LLM Wisdom
by: Li, Yisen, et al.
Published: (2025)
by: Li, Yisen, et al.
Published: (2025)
From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning
by: Li, Ming, et al.
Published: (2023)
by: Li, Ming, et al.
Published: (2023)
ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning
by: Wu, Yang, et al.
Published: (2024)
by: Wu, Yang, et al.
Published: (2024)
MADS: Model-Aware Diverse Core Set Selection for Instruction Tuning
by: Bai, Yi, et al.
Published: (2026)
by: Bai, Yi, et al.
Published: (2026)
IFIR: A Comprehensive Benchmark for Evaluating Instruction-Following in Expert-Domain Information Retrieval
by: Song, Tingyu, et al.
Published: (2025)
by: Song, Tingyu, et al.
Published: (2025)
Beyond Similarity: A Gradient-based Graph Method for Instruction Tuning Data Selection
by: Zhao, Yang, et al.
Published: (2025)
by: Zhao, Yang, et al.
Published: (2025)
Token Cleaning: Fine-Grained Data Selection for LLM Supervised Fine-Tuning
by: Pang, Jinlong, et al.
Published: (2025)
by: Pang, Jinlong, et al.
Published: (2025)
Rethinking Data Selection for Supervised Fine-Tuning
by: Shen, Ming
Published: (2024)
by: Shen, Ming
Published: (2024)
Data Selection for Multi-turn Dialogue Instruction Tuning
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
Retrieval Augmented Instruction Tuning for Open NER with Large Language Models
by: Xie, Tingyu, et al.
Published: (2024)
by: Xie, Tingyu, et al.
Published: (2024)
TACOS: Open Tagging and Comparative Scoring for Instruction Fine-Tuning Data Selection
by: He, Xixiang, et al.
Published: (2025)
by: He, Xixiang, et al.
Published: (2025)
From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM Alignment
by: Chen, Hao, et al.
Published: (2026)
by: Chen, Hao, et al.
Published: (2026)
InstructDiff: Domain-Adaptive Data Selection via Differential Entropy for Efficient LLM Fine-Tuning
by: Su, Junyou, et al.
Published: (2026)
by: Su, Junyou, et al.
Published: (2026)
SelectIT: Selective Instruction Tuning for LLMs via Uncertainty-Aware Self-Reflection
by: Liu, Liangxin, et al.
Published: (2024)
by: Liu, Liangxin, et al.
Published: (2024)
Instruction Mining: Instruction Data Selection for Tuning Large Language Models
by: Cao, Yihan, et al.
Published: (2023)
by: Cao, Yihan, et al.
Published: (2023)
Less is More: High-value Data Selection for Visual Instruction Tuning
by: Liu, Zikang, et al.
Published: (2024)
by: Liu, Zikang, et al.
Published: (2024)
Filter-then-Weight: Online Data Selection and Reweighting for LLM Fine-Tuning
by: Wang, Fangxin, et al.
Published: (2026)
by: Wang, Fangxin, et al.
Published: (2026)
Data Diversity Matters for Robust Instruction Tuning
by: Bukharin, Alexander, et al.
Published: (2023)
by: Bukharin, Alexander, et al.
Published: (2023)
LESS: Selecting Influential Data for Targeted Instruction Tuning
by: Xia, Mengzhou, et al.
Published: (2024)
by: Xia, Mengzhou, et al.
Published: (2024)
Similar Items
-
Co-EPG: A Framework for Co-Evolution of Planning and Grounding in Autonomous GUI Agents
by: Zhao, Yuan, et al.
Published: (2025) -
IterSelectTune: An Iterative Training Framework for Efficient Instruction-Tuning Data Selection
by: Song, Jielin, et al.
Published: (2024) -
CoachLM: Automatic Instruction Revisions Improve the Data Quality in LLM Instruction Tuning
by: Liu, Yilun, et al.
Published: (2023) -
A Survey on Data Selection for LLM Instruction Tuning
by: Zhang, Bolin, et al.
Published: (2024) -
Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning
by: Yang, Cehao, et al.
Published: (2025)