WikiHint: A Human-Annotated Dataset for Hint Ranking and Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mozafari, Jamshid, Gerhold, Florian, Jatowt, Adam
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913800652849152
author Mozafari, Jamshid
Gerhold, Florian
Jatowt, Adam
author_facet Mozafari, Jamshid
Gerhold, Florian
Jatowt, Adam
contents The use of Large Language Models (LLMs) has increased significantly with users frequently asking questions to chatbots. In the time when information is readily accessible, it is crucial to stimulate and preserve human cognitive abilities and maintain strong reasoning skills. This paper addresses such challenges by promoting the use of hints as an alternative or a supplement to direct answers. We first introduce a manually constructed hint dataset, WikiHint, which is based on Wikipedia and includes 5,000 hints created for 1,000 questions. We then finetune open-source LLMs for hint generation in answer-aware and answer-agnostic contexts. We assess the effectiveness of the hints with human participants who answer questions with and without the aid of hints. Additionally, we introduce a lightweight evaluation method, HintRank, to evaluate and rank hints in both answer-aware and answer-agnostic settings. Our findings show that (a) the dataset helps generate more effective hints, (b) including answer information along with questions generally improves the quality of generated hints, and (c) encoder-based models perform better than decoder-based models in hint ranking.
format Preprint
id arxiv_https___arxiv_org_abs_2412_01626
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle WikiHint: A Human-Annotated Dataset for Hint Ranking and Generation
Mozafari, Jamshid
Gerhold, Florian
Jatowt, Adam
Computation and Language
Information Retrieval
The use of Large Language Models (LLMs) has increased significantly with users frequently asking questions to chatbots. In the time when information is readily accessible, it is crucial to stimulate and preserve human cognitive abilities and maintain strong reasoning skills. This paper addresses such challenges by promoting the use of hints as an alternative or a supplement to direct answers. We first introduce a manually constructed hint dataset, WikiHint, which is based on Wikipedia and includes 5,000 hints created for 1,000 questions. We then finetune open-source LLMs for hint generation in answer-aware and answer-agnostic contexts. We assess the effectiveness of the hints with human participants who answer questions with and without the aid of hints. Additionally, we introduce a lightweight evaluation method, HintRank, to evaluate and rank hints in both answer-aware and answer-agnostic settings. Our findings show that (a) the dataset helps generate more effective hints, (b) including answer information along with questions generally improves the quality of generated hints, and (c) encoder-based models perform better than decoder-based models in hint ranking.
title WikiHint: A Human-Annotated Dataset for Hint Ranking and Generation
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2412.01626