FANNO: Augmenting High-Quality Instruction Data with Open-Sourced LLMs Only

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, He, Su, Junyou, Lun, Tianle, Tao, Yicheng, Zhang, Wenjia, Fan, Zipei, Chen, Guanhua
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910552577540096
author Zhu, He
Su, Junyou
Lun, Tianle
Tao, Yicheng
Zhang, Wenjia
Fan, Zipei
Chen, Guanhua
author_facet Zhu, He
Su, Junyou
Lun, Tianle
Tao, Yicheng
Zhang, Wenjia
Fan, Zipei
Chen, Guanhua
contents Instruction fine-tuning stands as a crucial advancement in leveraging large language models (LLMs) for enhanced task performance. However, the annotation of instruction datasets has traditionally been expensive and laborious, often relying on manual annotations or costly API calls of proprietary LLMs. To address these challenges, we introduce FANNO, a fully autonomous, open-sourced framework that revolutionizes the annotation process without the need for pre-existing annotated data. Utilizing a Mistral-7b-instruct model, FANNO efficiently produces diverse and high-quality datasets through a structured process involving document pre-screening, instruction generation, and response generation. Experiments on Open LLM Leaderboard and AlpacaEval benchmark show that the FANNO can generate high-quality data with diversity and complexity for free, comparable to human-annotated or cleaned datasets like Alpaca-GPT4-Cleaned.
format Preprint
id arxiv_https___arxiv_org_abs_2408_01323
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FANNO: Augmenting High-Quality Instruction Data with Open-Sourced LLMs Only
Zhu, He
Su, Junyou
Lun, Tianle
Tao, Yicheng
Zhang, Wenjia
Fan, Zipei
Chen, Guanhua
Computation and Language
Instruction fine-tuning stands as a crucial advancement in leveraging large language models (LLMs) for enhanced task performance. However, the annotation of instruction datasets has traditionally been expensive and laborious, often relying on manual annotations or costly API calls of proprietary LLMs. To address these challenges, we introduce FANNO, a fully autonomous, open-sourced framework that revolutionizes the annotation process without the need for pre-existing annotated data. Utilizing a Mistral-7b-instruct model, FANNO efficiently produces diverse and high-quality datasets through a structured process involving document pre-screening, instruction generation, and response generation. Experiments on Open LLM Leaderboard and AlpacaEval benchmark show that the FANNO can generate high-quality data with diversity and complexity for free, comparable to human-annotated or cleaned datasets like Alpaca-GPT4-Cleaned.
title FANNO: Augmenting High-Quality Instruction Data with Open-Sourced LLMs Only
topic Computation and Language
url https://arxiv.org/abs/2408.01323