The Right Time Matters: Data Arrangement Affects Zero-Shot Generalization in Instruction Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | He, Bingxiang, Ding, Ning, Qian, Cheng, Deng, Jia, Cui, Ganqu, Yuan, Lifan, Hong, Haiwen, Gao, Huan-ang, Huang, Longtao, Xue, Hui, Chen, Huimin, Liu, Zhiyuan, Sun, Maosong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AIR: A Systematic Analysis of Annotations, Instructions, and Response Pairs in Preference Dataset
by: He, Bingxiang, et al.
Published: (2025)
by: He, Bingxiang, et al.
Published: (2025)
UltraFeedback: Boosting Language Models with Scaled AI Feedback
by: Cui, Ganqu, et al.
Published: (2023)
by: Cui, Ganqu, et al.
Published: (2023)
Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
by: Guo, Yiju, et al.
Published: (2024)
by: Guo, Yiju, et al.
Published: (2024)
From $f(x)$ and $g(x)$ to $f(g(x))$: LLMs Learn New Skills in RL by Composing Old Ones
by: Yuan, Lifan, et al.
Published: (2025)
by: Yuan, Lifan, et al.
Published: (2025)
Advancing LLM Reasoning Generalists with Preference Trees
by: Yuan, Lifan, et al.
Published: (2024)
by: Yuan, Lifan, et al.
Published: (2024)
Free Process Rewards without Process Labels
by: Yuan, Lifan, et al.
Published: (2024)
by: Yuan, Lifan, et al.
Published: (2024)
RLPR: Extrapolating RLVR to General Domains without Verifiers
by: Yu, Tianyu, et al.
Published: (2025)
by: Yu, Tianyu, et al.
Published: (2025)
UniPSDA: Unsupervised Pseudo Semantic Data Augmentation for Zero-Shot Cross-Lingual Natural Language Understanding
by: Li, Dongyang, et al.
Published: (2024)
by: Li, Dongyang, et al.
Published: (2024)
Mastering Text, Code and Math Simultaneously via Fusing Highly Specialized Language Models
by: Ding, Ning, et al.
Published: (2024)
by: Ding, Ning, et al.
Published: (2024)
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
by: Li, Yaxuan, et al.
Published: (2026)
by: Li, Yaxuan, et al.
Published: (2026)
UltraIF: Advancing Instruction Following from the Wild
by: An, Kaikai, et al.
Published: (2025)
by: An, Kaikai, et al.
Published: (2025)
JustRL: Scaling a 1.5B LLM with a Simple RL Recipe
by: He, Bingxiang, et al.
Published: (2025)
by: He, Bingxiang, et al.
Published: (2025)
How LoRA Remembers? A Parametric Memory Law for LLM Finetuning
by: Xu, Ziwen, et al.
Published: (2026)
by: Xu, Ziwen, et al.
Published: (2026)
Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction Tuning
by: Ge, Chendi, et al.
Published: (2025)
by: Ge, Chendi, et al.
Published: (2025)
The Overthinker's DIET: Cutting Token Calories with DIfficulty-AwarE Training
by: Chen, Weize, et al.
Published: (2025)
by: Chen, Weize, et al.
Published: (2025)
Exploring Format Consistency for Instruction Tuning
by: Liang, Shihao, et al.
Published: (2023)
by: Liang, Shihao, et al.
Published: (2023)
INTERVENOR: Prompting the Coding Ability of Large Language Models with the Interactive Chain of Repair
by: Wang, Hanbin, et al.
Published: (2023)
by: Wang, Hanbin, et al.
Published: (2023)
Noise Contrastive Alignment of Language Models with Explicit Rewards
by: Chen, Huayu, et al.
Published: (2024)
by: Chen, Huayu, et al.
Published: (2024)
How Far Can Unsupervised RLVR Scale LLM Training?
by: He, Bingxiang, et al.
Published: (2026)
by: He, Bingxiang, et al.
Published: (2026)
Deep Exploration of Cross-Lingual Zero-Shot Generalization in Instruction Tuning
by: Han, Janghoon, et al.
Published: (2024)
by: Han, Janghoon, et al.
Published: (2024)
Learning to Generate Instruction Tuning Datasets for Zero-Shot Task Adaptation
by: Nayak, Nihal V., et al.
Published: (2024)
by: Nayak, Nihal V., et al.
Published: (2024)
Zero-Shot Code Representation Learning via Prompt Tuning
by: Cui, Nan, et al.
Published: (2024)
by: Cui, Nan, et al.
Published: (2024)
simpleposter: a simple baseline for product poster generation
by: Cui, Benlei, et al.
Published: (2026)
by: Cui, Benlei, et al.
Published: (2026)
FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
by: Hsieh, Cheng-Yu, et al.
Published: (2025)
by: Hsieh, Cheng-Yu, et al.
Published: (2025)
Zero-Shot Cross-Domain Code Search without Fine-Tuning
by: Liang, Keyu, et al.
Published: (2025)
by: Liang, Keyu, et al.
Published: (2025)
Tell Me More! Towards Implicit User Intention Understanding of Language Model Driven Agents
by: Qian, Cheng, et al.
Published: (2024)
by: Qian, Cheng, et al.
Published: (2024)
Unsupervised Text Representation Learning via Instruction-Tuning for Zero-Shot Dense Retrieval
by: Zeng, Qiuhai, et al.
Published: (2024)
by: Zeng, Qiuhai, et al.
Published: (2024)
Scalable Efficient Training of Large Language Models with Low-dimensional Projected Attention
by: Lv, Xingtai, et al.
Published: (2024)
by: Lv, Xingtai, et al.
Published: (2024)
Zero-Shot Scalable Resilience in UAV Swarms: A Decentralized Imitation Learning Framework with Physics-Informed Graph Interactions
by: Lin, Huan, et al.
Published: (2026)
by: Lin, Huan, et al.
Published: (2026)
Toward Zero-Shot Instruction Following
by: Lou, Renze, et al.
Published: (2023)
by: Lou, Renze, et al.
Published: (2023)
Dual-frame Fluid Motion Estimation with Test-time Optimization and Zero-divergence Loss
by: Zhang, Yifei, et al.
Published: (2024)
by: Zhang, Yifei, et al.
Published: (2024)
Diffusion Probe: Generated Image Result Prediction Using CNN Probes
by: Cui, Benlei, et al.
Published: (2026)
by: Cui, Benlei, et al.
Published: (2026)
Process Reinforcement through Implicit Rewards
by: Cui, Ganqu, et al.
Published: (2025)
by: Cui, Ganqu, et al.
Published: (2025)
In-Context Examples Matter: Improving Emotion Recognition in Conversation with Instruction Tuning
by: Ma, Hui, et al.
Published: (2025)
by: Ma, Hui, et al.
Published: (2025)
A representation of range decreasing group homomorphisms
by: Zhang, Ning, et al.
Published: (2025)
by: Zhang, Ning, et al.
Published: (2025)
Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
by: Wang, Xing, et al.
Published: (2025)
by: Wang, Xing, et al.
Published: (2025)
Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages
by: Hu, Jinyi, et al.
Published: (2023)
by: Hu, Jinyi, et al.
Published: (2023)
Coherent Zero-Shot Visual Instruction Generation
by: Phung, Quynh, et al.
Published: (2024)
by: Phung, Quynh, et al.
Published: (2024)
Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model in End-to-End Speech Recognition
by: Higuchi, Yosuke, et al.
Published: (2023)
by: Higuchi, Yosuke, et al.
Published: (2023)
Right-Side-Out: Learning Zero-Shot Sim-to-Real Garment Reversal
by: Yu, Chang, et al.
Published: (2025)
by: Yu, Chang, et al.
Published: (2025)
Similar Items
-
AIR: A Systematic Analysis of Annotations, Instructions, and Response Pairs in Preference Dataset
by: He, Bingxiang, et al.
Published: (2025) -
UltraFeedback: Boosting Language Models with Scaled AI Feedback
by: Cui, Ganqu, et al.
Published: (2023) -
Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
by: Guo, Yiju, et al.
Published: (2024) -
From $f(x)$ and $g(x)$ to $f(g(x))$: LLMs Learn New Skills in RL by Composing Old Ones
by: Yuan, Lifan, et al.
Published: (2025) -
Advancing LLM Reasoning Generalists with Preference Trees
by: Yuan, Lifan, et al.
Published: (2024)