What Makes Good Instruction-Tuning Data? An In-Context Learning Perspective
Fuente:
arXiv
Salvato in:
| Autori principali: | Han, Guangzeng, Huang, Xiaolei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Attributes as Textual Genes: Leveraging LLMs as Genetic Algorithm Simulators for Conditional Synthetic Data Generation
di: Han, Guangzeng, et al.
Pubblicazione: (2025)
di: Han, Guangzeng, et al.
Pubblicazione: (2025)
Model-Agnostic Meta Learning for Class Imbalance Adaptation
di: Rao, Hanshu, et al.
Pubblicazione: (2026)
di: Rao, Hanshu, et al.
Pubblicazione: (2026)
Chain-of-Interaction: Enhancing Large Language Models for Psychiatric Behavior Understanding by Dyadic Contexts
di: Han, Guangzeng, et al.
Pubblicazione: (2024)
di: Han, Guangzeng, et al.
Pubblicazione: (2024)
Knowledge-driven Augmentation and Retrieval for Integrative Temporal Adaptation
di: Liu, Weisi, et al.
Pubblicazione: (2026)
di: Liu, Weisi, et al.
Pubblicazione: (2026)
Length-Aware Multi-Kernel Transformer for Long Document Classification
di: Han, Guangzeng, et al.
Pubblicazione: (2024)
di: Han, Guangzeng, et al.
Pubblicazione: (2024)
Examining and Adapting Time for Multilingual Classification via Mixture of Temporal Experts
di: Liu, Weisi, et al.
Pubblicazione: (2025)
di: Liu, Weisi, et al.
Pubblicazione: (2025)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
di: Du, Yifan, et al.
Pubblicazione: (2023)
di: Du, Yifan, et al.
Pubblicazione: (2023)
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
di: Liu, Wei, et al.
Pubblicazione: (2023)
di: Liu, Wei, et al.
Pubblicazione: (2023)
Sample Design Engineering: An Empirical Study of What Makes Good Downstream Fine-Tuning Samples for LLMs
di: Guo, Biyang, et al.
Pubblicazione: (2024)
di: Guo, Biyang, et al.
Pubblicazione: (2024)
Leveraging Multimodal Self-Consistency Reasoning in Coding Motivational Interviewing for Alcohol Use Reduction
di: Han, Guangzeng, et al.
Pubblicazione: (2026)
di: Han, Guangzeng, et al.
Pubblicazione: (2026)
How Many Languages Make Good Multilingual Instruction Tuning? A Case Study on BLOOM
di: Ji, Shaoxiong, et al.
Pubblicazione: (2024)
di: Ji, Shaoxiong, et al.
Pubblicazione: (2024)
What Makes Language Models Good-enough?
di: Asami, Daiki, et al.
Pubblicazione: (2024)
di: Asami, Daiki, et al.
Pubblicazione: (2024)
What Makes a Good Natural Language Prompt?
di: Long, Do Xuan, et al.
Pubblicazione: (2025)
di: Long, Do Xuan, et al.
Pubblicazione: (2025)
RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection
di: Yang, Yixin, et al.
Pubblicazione: (2025)
di: Yang, Yixin, et al.
Pubblicazione: (2025)
What Makes a Reward Model a Good Teacher? An Optimization Perspective
di: Razin, Noam, et al.
Pubblicazione: (2025)
di: Razin, Noam, et al.
Pubblicazione: (2025)
Does Instruction Tuning Make LLMs More Consistent?
di: Fierro, Constanza, et al.
Pubblicazione: (2024)
di: Fierro, Constanza, et al.
Pubblicazione: (2024)
Generalizing From Short to Long: Effective Data Synthesis for Long-Context Instruction Tuning
di: Zhu, Wenhao, et al.
Pubblicazione: (2025)
di: Zhu, Wenhao, et al.
Pubblicazione: (2025)
Is It Good Data for Multilingual Instruction Tuning or Just Bad Multilingual Evaluation for Large Language Models?
di: Chen, Pinzhen, et al.
Pubblicazione: (2024)
di: Chen, Pinzhen, et al.
Pubblicazione: (2024)
Not All Documents Are What You Need for Extracting Instruction Tuning Data
di: Zhang, Chi, et al.
Pubblicazione: (2025)
di: Zhang, Chi, et al.
Pubblicazione: (2025)
The Inherent Limits of Pretrained LLMs: The Unexpected Convergence of Instruction Tuning and In-Context Learning Capabilities
di: Bigoulaeva, Irina, et al.
Pubblicazione: (2025)
di: Bigoulaeva, Irina, et al.
Pubblicazione: (2025)
PACIT: Unlocking the Power of Examples for Better In-Context Instruction Tuning
di: Xue, Tianci, et al.
Pubblicazione: (2023)
di: Xue, Tianci, et al.
Pubblicazione: (2023)
Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages
di: Toukmaji, Christopher, et al.
Pubblicazione: (2025)
di: Toukmaji, Christopher, et al.
Pubblicazione: (2025)
Hint Tuning: Less Data Makes Better Reasoners
di: Fan, Siqi, et al.
Pubblicazione: (2026)
di: Fan, Siqi, et al.
Pubblicazione: (2026)
What Makes a Good Doctor Response? A Study on Text-Based Telemedicine
di: Cosma, Adrian, et al.
Pubblicazione: (2026)
di: Cosma, Adrian, et al.
Pubblicazione: (2026)
DoG-Instruct: Towards Premium Instruction-Tuning Data via Text-Grounded Instruction Wrapping
di: Chen, Yongrui, et al.
Pubblicazione: (2023)
di: Chen, Yongrui, et al.
Pubblicazione: (2023)
A Scoping Review of Synthetic Data Generation by Language Models in Biomedical Research and Application: Data Utility and Quality Perspectives
di: Rao, Hanshu, et al.
Pubblicazione: (2025)
di: Rao, Hanshu, et al.
Pubblicazione: (2025)
In-Context Examples Matter: Improving Emotion Recognition in Conversation with Instruction Tuning
di: Ma, Hui, et al.
Pubblicazione: (2025)
di: Ma, Hui, et al.
Pubblicazione: (2025)
Large Language Models Know What Makes Exemplary Contexts
di: Long, Quanyu, et al.
Pubblicazione: (2024)
di: Long, Quanyu, et al.
Pubblicazione: (2024)
Large-Scale Data Selection for Instruction Tuning
di: Ivison, Hamish, et al.
Pubblicazione: (2025)
di: Ivison, Hamish, et al.
Pubblicazione: (2025)
Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models
di: Li, Haoran, et al.
Pubblicazione: (2024)
di: Li, Haoran, et al.
Pubblicazione: (2024)
What Makes a Good Response? An Empirical Analysis of Quality in Qualitative Interviews
di: Ivey, Jonathan, et al.
Pubblicazione: (2026)
di: Ivey, Jonathan, et al.
Pubblicazione: (2026)
What Makes Good Multilingual Reasoning? Disentangling Reasoning Traces with Measurable Features
di: Ki, Dayeon, et al.
Pubblicazione: (2026)
di: Ki, Dayeon, et al.
Pubblicazione: (2026)
Scaling Instruction-Tuned LLMs to Million-Token Contexts via Hierarchical Synthetic Data Generation
di: He, Linda, et al.
Pubblicazione: (2025)
di: He, Linda, et al.
Pubblicazione: (2025)
Understanding LLMs' Cross-Lingual Context Retrieval: How Good It Is And Where It Comes From
di: Gao, Changjiang, et al.
Pubblicazione: (2025)
di: Gao, Changjiang, et al.
Pubblicazione: (2025)
DS$^2$-Instruct: Domain-Specific Data Synthesis for Large Language Models Instruction Tuning
di: Xu, Ruiyao, et al.
Pubblicazione: (2026)
di: Xu, Ruiyao, et al.
Pubblicazione: (2026)
Otter: A Multi-Modal Model with In-Context Instruction Tuning
di: Li, Bo, et al.
Pubblicazione: (2023)
di: Li, Bo, et al.
Pubblicazione: (2023)
Neuron-Aware Data Selection In Instruction Tuning For Large Language Models
di: Chen, Xin, et al.
Pubblicazione: (2026)
di: Chen, Xin, et al.
Pubblicazione: (2026)
From UAV Imagery to Agronomic Reasoning: A Multimodal LLM Benchmark for Plant Phenotyping
di: Wu, Yu, et al.
Pubblicazione: (2026)
di: Wu, Yu, et al.
Pubblicazione: (2026)
Cultivating Multidisciplinary AI Workforce Development on iTiger GPU Cluster: Practices and Challenges
di: Sharif, Mayira, et al.
Pubblicazione: (2025)
di: Sharif, Mayira, et al.
Pubblicazione: (2025)
Instruction Tuning with and without Context: Behavioral Shifts and Downstream Impact
di: Lee, Hyunji, et al.
Pubblicazione: (2025)
di: Lee, Hyunji, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Attributes as Textual Genes: Leveraging LLMs as Genetic Algorithm Simulators for Conditional Synthetic Data Generation
di: Han, Guangzeng, et al.
Pubblicazione: (2025) -
Model-Agnostic Meta Learning for Class Imbalance Adaptation
di: Rao, Hanshu, et al.
Pubblicazione: (2026) -
Chain-of-Interaction: Enhancing Large Language Models for Psychiatric Behavior Understanding by Dyadic Contexts
di: Han, Guangzeng, et al.
Pubblicazione: (2024) -
Knowledge-driven Augmentation and Retrieval for Integrative Temporal Adaptation
di: Liu, Weisi, et al.
Pubblicazione: (2026) -
Length-Aware Multi-Kernel Transformer for Long Document Classification
di: Han, Guangzeng, et al.
Pubblicazione: (2024)