MAIN: Mutual Alignment Is Necessary for instruction tuning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Fanyi, Liu, Jianfeng, Zhang, Xin, Liu, Haoyu, Cao, Xixin, Zhan, Yuefeng, Sun, Hao, Deng, Weiwei, Sun, Feng, Zhang, Qi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909967995371520
author Yang, Fanyi
Liu, Jianfeng
Zhang, Xin
Liu, Haoyu
Cao, Xixin
Zhan, Yuefeng
Sun, Hao
Deng, Weiwei
Sun, Feng
Zhang, Qi
author_facet Yang, Fanyi
Liu, Jianfeng
Zhang, Xin
Liu, Haoyu
Cao, Xixin
Zhan, Yuefeng
Sun, Hao
Deng, Weiwei
Sun, Feng
Zhang, Qi
contents Instruction tuning has empowered large language models (LLMs) to achieve remarkable performance, yet its success heavily depends on the availability of large-scale, high-quality instruction-response pairs. To meet this demand, various methods have been developed to synthesize data at scale. However, current methods for scaling up data generation often overlook a crucial aspect: the alignment between instructions and responses. We hypothesize that the quality of instruction-response pairs is determined not by the individual quality of each component, but by the degree of mutual alignment. To address this, we propose a Mutual Alignment Framework (MAIN) which enforces coherence between instructions and responses through mutual constraints. We demonstrate that MAIN generalizes well across model architectures and sizes, achieving state-of-the-art performance on LLaMA, Mistral, and Qwen models across diverse benchmarks. This work underscores the critical role of instruction-response alignment in enabling generalizable and high-quality instruction tuning for LLMs. All code is available from our repository.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12913
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MAIN: Mutual Alignment Is Necessary for instruction tuning
Yang, Fanyi
Liu, Jianfeng
Zhang, Xin
Liu, Haoyu
Cao, Xixin
Zhan, Yuefeng
Sun, Hao
Deng, Weiwei
Sun, Feng
Zhang, Qi
Computation and Language
Instruction tuning has empowered large language models (LLMs) to achieve remarkable performance, yet its success heavily depends on the availability of large-scale, high-quality instruction-response pairs. To meet this demand, various methods have been developed to synthesize data at scale. However, current methods for scaling up data generation often overlook a crucial aspect: the alignment between instructions and responses. We hypothesize that the quality of instruction-response pairs is determined not by the individual quality of each component, but by the degree of mutual alignment. To address this, we propose a Mutual Alignment Framework (MAIN) which enforces coherence between instructions and responses through mutual constraints. We demonstrate that MAIN generalizes well across model architectures and sizes, achieving state-of-the-art performance on LLaMA, Mistral, and Qwen models across diverse benchmarks. This work underscores the critical role of instruction-response alignment in enabling generalizable and high-quality instruction tuning for LLMs. All code is available from our repository.
title MAIN: Mutual Alignment Is Necessary for instruction tuning
topic Computation and Language
url https://arxiv.org/abs/2504.12913