InFoBench: Evaluating Instruction Following Ability in Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Qin, Yiwei, Song, Kaiqiang, Hu, Yebowen, Yao, Wenlin, Cho, Sangwoo, Wang, Xiaoyang, Wu, Xuansheng, Liu, Fei, Liu, Pengfei, Yu, Dong
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916083779239936
author Qin, Yiwei
Song, Kaiqiang
Hu, Yebowen
Yao, Wenlin
Cho, Sangwoo
Wang, Xiaoyang
Wu, Xuansheng
Liu, Fei
Liu, Pengfei
Yu, Dong
author_facet Qin, Yiwei
Song, Kaiqiang
Hu, Yebowen
Yao, Wenlin
Cho, Sangwoo
Wang, Xiaoyang
Wu, Xuansheng
Liu, Fei
Liu, Pengfei
Yu, Dong
contents This paper introduces the Decomposed Requirements Following Ratio (DRFR), a new metric for evaluating Large Language Models' (LLMs) ability to follow instructions. Addressing a gap in current methodologies, DRFR breaks down complex instructions into simpler criteria, facilitating a detailed analysis of LLMs' compliance with various aspects of tasks. Alongside this metric, we present InFoBench, a benchmark comprising 500 diverse instructions and 2,250 decomposed questions across multiple constraint categories. Our experiments compare DRFR with traditional scoring methods and explore annotation sources, including human experts, crowd-sourced workers, and GPT-4. The findings demonstrate DRFR's higher reliability and the effectiveness of using GPT-4 as a cost-efficient annotator. The evaluation of several advanced LLMs using this framework reveals their strengths and areas needing improvement, particularly in complex instruction-following. This study contributes a novel metric and benchmark, offering insights for future LLM development and evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2401_03601
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle InFoBench: Evaluating Instruction Following Ability in Large Language Models
Qin, Yiwei
Song, Kaiqiang
Hu, Yebowen
Yao, Wenlin
Cho, Sangwoo
Wang, Xiaoyang
Wu, Xuansheng
Liu, Fei
Liu, Pengfei
Yu, Dong
Computation and Language
Artificial Intelligence
This paper introduces the Decomposed Requirements Following Ratio (DRFR), a new metric for evaluating Large Language Models' (LLMs) ability to follow instructions. Addressing a gap in current methodologies, DRFR breaks down complex instructions into simpler criteria, facilitating a detailed analysis of LLMs' compliance with various aspects of tasks. Alongside this metric, we present InFoBench, a benchmark comprising 500 diverse instructions and 2,250 decomposed questions across multiple constraint categories. Our experiments compare DRFR with traditional scoring methods and explore annotation sources, including human experts, crowd-sourced workers, and GPT-4. The findings demonstrate DRFR's higher reliability and the effectiveness of using GPT-4 as a cost-efficient annotator. The evaluation of several advanced LLMs using this framework reveals their strengths and areas needing improvement, particularly in complex instruction-following. This study contributes a novel metric and benchmark, offering insights for future LLM development and evaluation.
title InFoBench: Evaluating Instruction Following Ability in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2401.03601