SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | He, Yexiao, Wang, Ziyao, Shen, Zheyu, Sun, Guoheng, Dai, Yucong, Wu, Yongkai, Wang, Hongyi, Li, Ang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations
by: Wang, Ziyao, et al.
Published: (2024)
by: Wang, Ziyao, et al.
Published: (2024)
Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation
by: Wang, Ziyao, et al.
Published: (2025)
by: Wang, Ziyao, et al.
Published: (2025)
Revisiting Federated Fine-Tuning: A Single Communication Round is Enough for Foundation Models
by: Wang, Ziyao, et al.
Published: (2024)
by: Wang, Ziyao, et al.
Published: (2024)
FedMOA: Federated GRPO for Personalized Reasoning LLMs under Heterogeneous Rewards
by: Wang, Ziyao, et al.
Published: (2026)
by: Wang, Ziyao, et al.
Published: (2026)
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
by: Shen, Zheyu, et al.
Published: (2025)
by: Shen, Zheyu, et al.
Published: (2025)
Prada: Black-Box LLM Adaptation with Private Data on Resource-Constrained Devices
by: Wang, Ziyao, et al.
Published: (2025)
by: Wang, Ziyao, et al.
Published: (2025)
What Matters in Transformers? Not All Attention is Needed
by: He, Shwai, et al.
Published: (2024)
by: He, Shwai, et al.
Published: (2024)
Fair Diagnosis: Leveraging Causal Modeling to Mitigate Medical Bias
by: Tian, Bowei, et al.
Published: (2024)
by: Tian, Bowei, et al.
Published: (2024)
Towards counterfactual fairness through auxiliary variables
by: Tian, Bowei, et al.
Published: (2024)
by: Tian, Bowei, et al.
Published: (2024)
SymRTLO: Enhancing RTL Code Optimization with LLMs and Neuron-Inspired Symbolic Reasoning
by: Wang, Yiting, et al.
Published: (2025)
by: Wang, Yiting, et al.
Published: (2025)
Invisible Tokens, Visible Bills: The Urgent Need to Audit Hidden Operations in Opaque LLM Services
by: Sun, Guoheng, et al.
Published: (2025)
by: Sun, Guoheng, et al.
Published: (2025)
Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines
by: Wang, Ziyao, et al.
Published: (2026)
by: Wang, Ziyao, et al.
Published: (2026)
Towards Building Non-Fine-Tunable Foundation Models
by: Wang, Ziyao, et al.
Published: (2026)
by: Wang, Ziyao, et al.
Published: (2026)
ROCKET: Residual-Oriented Multi-Layer Alignment for Spatially-Aware Vision-Language-Action Models
by: Sun, Guoheng, et al.
Published: (2026)
by: Sun, Guoheng, et al.
Published: (2026)
Demystifying When Pruning Works via Representation Hierarchies
by: He, Shwai, et al.
Published: (2026)
by: He, Shwai, et al.
Published: (2026)
Model-GLUE: Democratized LLM Scaling for A Large Model Zoo in the Wild
by: Zhao, Xinyu, et al.
Published: (2024)
by: Zhao, Xinyu, et al.
Published: (2024)
CoIn: Counting the Invisible Reasoning Tokens in Commercial Opaque LLM APIs
by: Sun, Guoheng, et al.
Published: (2025)
by: Sun, Guoheng, et al.
Published: (2025)
Diversity Measurement and Subset Selection for Instruction Tuning Datasets
by: Wang, Peiqi, et al.
Published: (2024)
by: Wang, Peiqi, et al.
Published: (2024)
VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
by: Wang, Yiting, et al.
Published: (2025)
by: Wang, Yiting, et al.
Published: (2025)
MindCraft: How Concept Trees Take Shape In Deep Models
by: Tian, Bowei, et al.
Published: (2025)
by: Tian, Bowei, et al.
Published: (2025)
FairSAM: Fair Classification on Corrupted Data Through Sharpness-Aware Minimization
by: Dai, Yucong, et al.
Published: (2025)
by: Dai, Yucong, et al.
Published: (2025)
Stabilizing LLM Supervised Fine-Tuning via Explicit Distributional Control
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
Augmenting NER Datasets with LLMs: Towards Automated and Refined Annotation
by: Naraki, Yuji, et al.
Published: (2024)
by: Naraki, Yuji, et al.
Published: (2024)
Proximal Supervised Fine-Tuning
by: Zhu, Wenhong, et al.
Published: (2025)
by: Zhu, Wenhong, et al.
Published: (2025)
Instruction Mining: Instruction Data Selection for Tuning Large Language Models
by: Cao, Yihan, et al.
Published: (2023)
by: Cao, Yihan, et al.
Published: (2023)
TokenShapley: Token Level Context Attribution with Shapley Value
by: Xiao, Yingtai, et al.
Published: (2025)
by: Xiao, Yingtai, et al.
Published: (2025)
Instruction Fine-Tuning: Does Prompt Loss Matter?
by: Huerta-Enochian, Mathew, et al.
Published: (2024)
by: Huerta-Enochian, Mathew, et al.
Published: (2024)
Low-Confidence Gold: Refining Low-Confidence Samples for Efficient Instruction Tuning
by: Cai, Hongyi, et al.
Published: (2025)
by: Cai, Hongyi, et al.
Published: (2025)
Deeper Insights Without Updates: The Power of In-Context Learning Over Fine-Tuning
by: Yin, Qingyu, et al.
Published: (2024)
by: Yin, Qingyu, et al.
Published: (2024)
Learning to Generate Instruction Tuning Datasets for Zero-Shot Task Adaptation
by: Nayak, Nihal V., et al.
Published: (2024)
by: Nayak, Nihal V., et al.
Published: (2024)
MedOrch: Medical Diagnosis with Tool-Augmented Reasoning Agents for Flexible Extensibility
by: He, Yexiao, et al.
Published: (2025)
by: He, Yexiao, et al.
Published: (2025)
The Best Instruction-Tuning Data are Those That Fit
by: Zhang, Dylan, et al.
Published: (2025)
by: Zhang, Dylan, et al.
Published: (2025)
Instruction Tuning for Large Language Models: A Survey
by: Zhang, Shengyu, et al.
Published: (2023)
by: Zhang, Shengyu, et al.
Published: (2023)
CoMMIT: Coordinated Multimodal Instruction Tuning
by: Li, Xintong, et al.
Published: (2024)
by: Li, Xintong, et al.
Published: (2024)
Parameter Efficient Instruction Tuning: An Empirical Study
by: He, Pengfei
Published: (2024)
by: He, Pengfei
Published: (2024)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
by: Hui, Tingfeng, et al.
Published: (2024)
by: Hui, Tingfeng, et al.
Published: (2024)
Data Shapley in One Training Run
by: Wang, Jiachen T., et al.
Published: (2024)
by: Wang, Jiachen T., et al.
Published: (2024)
Anchored Supervised Fine-Tuning
by: Zhu, He, et al.
Published: (2025)
by: Zhu, He, et al.
Published: (2025)
Federated Data-Efficient Instruction Tuning for Large Language Models
by: Qin, Zhen, et al.
Published: (2024)
by: Qin, Zhen, et al.
Published: (2024)
FairAgent: Democratizing Fairness-Aware Machine Learning with LLM-Powered Agents
by: Dai, Yucong, et al.
Published: (2025)
by: Dai, Yucong, et al.
Published: (2025)
Similar Items
-
FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations
by: Wang, Ziyao, et al.
Published: (2024) -
Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation
by: Wang, Ziyao, et al.
Published: (2025) -
Revisiting Federated Fine-Tuning: A Single Communication Round is Enough for Foundation Models
by: Wang, Ziyao, et al.
Published: (2024) -
FedMOA: Federated GRPO for Personalized Reasoning LLMs under Heterogeneous Rewards
by: Wang, Ziyao, et al.
Published: (2026) -
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
by: Shen, Zheyu, et al.
Published: (2025)