SelectIT: Selective Instruction Tuning for LLMs via Uncertainty-Aware Self-Reflection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Liangxin, Liu, Xuebo, Wong, Derek F., Li, Dongfang, Wang, Ziyi, Hu, Baotian, Zhang, Min |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LycheeCluster: Efficient Long-Context Inference with Structure-Aware Chunking and Hierarchical KV Indexing
von: Li, Dongfang, et al.
Veröffentlicht: (2026)
von: Li, Dongfang, et al.
Veröffentlicht: (2026)
Improving Attributed Text Generation of Large Language Models via Preference Learning
von: Li, Dongfang, et al.
Veröffentlicht: (2024)
von: Li, Dongfang, et al.
Veröffentlicht: (2024)
Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
Uncertainty-Aware Gradient Signal-to-Noise Data Selection for Instruction Tuning
von: Yuan, Zhihang, et al.
Veröffentlicht: (2026)
von: Yuan, Zhihang, et al.
Veröffentlicht: (2026)
Improving Value-based Process Verifier via Structural Prior Injection
von: Sun, Zetian, et al.
Veröffentlicht: (2025)
von: Sun, Zetian, et al.
Veröffentlicht: (2025)
CommonIT: Commonality-Aware Instruction Tuning for Large Language Models via Data Partitions
von: Rao, Jun, et al.
Veröffentlicht: (2024)
von: Rao, Jun, et al.
Veröffentlicht: (2024)
Improving Value-based Process Verifier via Low-Cost Variance Reduction
von: Sun, Zetian, et al.
Veröffentlicht: (2025)
von: Sun, Zetian, et al.
Veröffentlicht: (2025)
Selective Prompting Tuning for Personalized Conversations with LLMs
von: Huang, Qiushi, et al.
Veröffentlicht: (2024)
von: Huang, Qiushi, et al.
Veröffentlicht: (2024)
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training
von: Liang, Yu, et al.
Veröffentlicht: (2026)
von: Liang, Yu, et al.
Veröffentlicht: (2026)
Stabilizing Long-term Multi-turn Reinforcement Learning with Gated Rewards
von: Sun, Zetian, et al.
Veröffentlicht: (2025)
von: Sun, Zetian, et al.
Veröffentlicht: (2025)
Loong: A Human-Like Long Document Translation Agent with Observe-and-Act Adaptive Context Selection
von: Wang, Yutong, et al.
Veröffentlicht: (2026)
von: Wang, Yutong, et al.
Veröffentlicht: (2026)
Dynamic Long Context Reasoning over Compressed Memory via End-to-End Reinforcement Learning
von: Chen, Zhuoen, et al.
Veröffentlicht: (2026)
von: Chen, Zhuoen, et al.
Veröffentlicht: (2026)
Instruction Mining: Instruction Data Selection for Tuning Large Language Models
von: Cao, Yihan, et al.
Veröffentlicht: (2023)
von: Cao, Yihan, et al.
Veröffentlicht: (2023)
LESS: Selecting Influential Data for Targeted Instruction Tuning
von: Xia, Mengzhou, et al.
Veröffentlicht: (2024)
von: Xia, Mengzhou, et al.
Veröffentlicht: (2024)
In-Context Learning State Vector with Inner and Momentum Optimization
von: Li, Dongfang, et al.
Veröffentlicht: (2024)
von: Li, Dongfang, et al.
Veröffentlicht: (2024)
Mitigating Training Imbalance in LLM Fine-Tuning via Selective Parameter Merging
von: Ju, Yiming, et al.
Veröffentlicht: (2024)
von: Ju, Yiming, et al.
Veröffentlicht: (2024)
CMT: A Memory Compression Method for Continual Knowledge Learning of Large Language Models
von: Li, Dongfang, et al.
Veröffentlicht: (2024)
von: Li, Dongfang, et al.
Veröffentlicht: (2024)
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
von: Liu, Wei, et al.
Veröffentlicht: (2023)
von: Liu, Wei, et al.
Veröffentlicht: (2023)
ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning
von: Wu, Yang, et al.
Veröffentlicht: (2024)
von: Wu, Yang, et al.
Veröffentlicht: (2024)
Is On-Policy Data always the Best Choice for Direct Preference Optimization-based LM Alignment?
von: Sun, Zetian, et al.
Veröffentlicht: (2025)
von: Sun, Zetian, et al.
Veröffentlicht: (2025)
Temporal Knowledge Question Answering via Abstract Reasoning Induction
von: Chen, Ziyang, et al.
Veröffentlicht: (2023)
von: Chen, Ziyang, et al.
Veröffentlicht: (2023)
TAGCOS: Task-agnostic Gradient Clustered Coreset Selection for Instruction Tuning Data
von: Zhang, Jipeng, et al.
Veröffentlicht: (2024)
von: Zhang, Jipeng, et al.
Veröffentlicht: (2024)
ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework
von: Qin, Kai, et al.
Veröffentlicht: (2026)
von: Qin, Kai, et al.
Veröffentlicht: (2026)
DIVE: Embedding Compression via Self-Limiting Gradient Updates
von: Zhao, Dongfang
Veröffentlicht: (2026)
von: Zhao, Dongfang
Veröffentlicht: (2026)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
von: Wu, Xuansheng, et al.
Veröffentlicht: (2023)
von: Wu, Xuansheng, et al.
Veröffentlicht: (2023)
M$^2$PT: Multimodal Prompt Tuning for Zero-shot Instruction Learning
von: Wang, Taowen, et al.
Veröffentlicht: (2024)
von: Wang, Taowen, et al.
Veröffentlicht: (2024)
Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking
von: Xu, Qinwu, et al.
Veröffentlicht: (2026)
von: Xu, Qinwu, et al.
Veröffentlicht: (2026)
TasTe: Teaching Large Language Models to Translate through Self-Reflection
von: Wang, Yutong, et al.
Veröffentlicht: (2024)
von: Wang, Yutong, et al.
Veröffentlicht: (2024)
SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales
von: Xu, Tianyang, et al.
Veröffentlicht: (2024)
von: Xu, Tianyang, et al.
Veröffentlicht: (2024)
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
von: Lin, Gang, et al.
Veröffentlicht: (2026)
von: Lin, Gang, et al.
Veröffentlicht: (2026)
Instruction Tuning with Human Curriculum
von: Lee, Bruce W., et al.
Veröffentlicht: (2023)
von: Lee, Bruce W., et al.
Veröffentlicht: (2023)
Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
ClusterUCB: Efficient Gradient-Based Data Selection for Targeted Fine-Tuning of LLMs
von: Wang, Zige, et al.
Veröffentlicht: (2025)
von: Wang, Zige, et al.
Veröffentlicht: (2025)
Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMs
von: Kang, Feiyang, et al.
Veröffentlicht: (2024)
von: Kang, Feiyang, et al.
Veröffentlicht: (2024)
Compositional Literary Primitives in Instruction-Tuned LLMs: Cross-Architectural SAE Features for Self, Style, and Affect
von: Presa, Joao Paulo Cavalcante, et al.
Veröffentlicht: (2026)
von: Presa, Joao Paulo Cavalcante, et al.
Veröffentlicht: (2026)
CYCLE-INSTRUCT: Fully Seed-Free Instruction Tuning via Dual Self-Training and Cycle Consistency
von: Shen, Zhanming, et al.
Veröffentlicht: (2025)
von: Shen, Zhanming, et al.
Veröffentlicht: (2025)
Selection of LLM Fine-Tuning Data based on Orthogonal Rules
von: Li, Xiaomin, et al.
Veröffentlicht: (2024)
von: Li, Xiaomin, et al.
Veröffentlicht: (2024)
Training-Trajectory-Aware Token Selection
von: Shen, Zhanming, et al.
Veröffentlicht: (2026)
von: Shen, Zhanming, et al.
Veröffentlicht: (2026)
Building Accurate Translation-Tailored LLMs with Language Aware Instruction Tuning
von: Zan, Changtong, et al.
Veröffentlicht: (2024)
von: Zan, Changtong, et al.
Veröffentlicht: (2024)
SelectLLM: Can LLMs Select Important Instructions to Annotate?
von: Parkar, Ritik Sachin, et al.
Veröffentlicht: (2024)
von: Parkar, Ritik Sachin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LycheeCluster: Efficient Long-Context Inference with Structure-Aware Chunking and Hierarchical KV Indexing
von: Li, Dongfang, et al.
Veröffentlicht: (2026) -
Improving Attributed Text Generation of Large Language Models via Preference Learning
von: Li, Dongfang, et al.
Veröffentlicht: (2024) -
Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning
von: Li, Ming, et al.
Veröffentlicht: (2024) -
Uncertainty-Aware Gradient Signal-to-Noise Data Selection for Instruction Tuning
von: Yuan, Zhihang, et al.
Veröffentlicht: (2026) -
Improving Value-based Process Verifier via Structural Prior Injection
von: Sun, Zetian, et al.
Veröffentlicht: (2025)