KNIGHT: Knowledge Graph-Driven Multiple-Choice Question Generation with Adaptive Hardness Calibration

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Amanlou, Mohammad, Moghaddam, Erfan Shafiee, Jafari, Yasaman Amou, Noori, Mahdi, Farsi, Farhan, Bahrak, Behnam
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918351286042624
author Amanlou, Mohammad
Moghaddam, Erfan Shafiee
Jafari, Yasaman Amou
Noori, Mahdi
Farsi, Farhan
Bahrak, Behnam
author_facet Amanlou, Mohammad
Moghaddam, Erfan Shafiee
Jafari, Yasaman Amou
Noori, Mahdi
Farsi, Farhan
Bahrak, Behnam
contents With the rise of large language models (LLMs), they have become instrumental in applications such as Retrieval-Augmented Generation (RAG). Yet evaluating these systems remains bottlenecked by the time and cost of building specialized assessment datasets. We introduce KNIGHT, an LLM-based, knowledge-graph-driven framework for generating multiple-choice question (MCQ) datasets from external sources. KNIGHT constructs a topic-specific knowledge graph, a structured and parsimonious summary of entities and relations, that can be reused to generate instructor-controlled difficulty levels, including multi-hop questions, without repeatedly re-feeding the full source text. This knowledge graph acts as a compressed, reusable state, making question generation a cheap read over the graph. We instantiate KNIGHT on Wikipedia/Wikidata while keeping the framework domain- and ontology-agnostic. As a case study, KNIGHT produces six MCQ datasets in History, Biology, and Mathematics. We evaluate quality on five criteria: fluency, unambiguity (single correct answer), topic relevance, option uniqueness, and answerability given the provided sources (as a proxy for hallucination). Results show that KNIGHT enables token- and cost-efficient generation from a reusable graph representation, achieves high quality across these criteria, and yields model rankings aligned with MMLU-style benchmarks, while supporting topic-specific and difficulty-controlled evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2602_20135
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle KNIGHT: Knowledge Graph-Driven Multiple-Choice Question Generation with Adaptive Hardness Calibration
Amanlou, Mohammad
Moghaddam, Erfan Shafiee
Jafari, Yasaman Amou
Noori, Mahdi
Farsi, Farhan
Bahrak, Behnam
Computation and Language
Artificial Intelligence
Information Retrieval
68T50, 68T05, 68P20
I.2.7; I.2.4; H.3.3
With the rise of large language models (LLMs), they have become instrumental in applications such as Retrieval-Augmented Generation (RAG). Yet evaluating these systems remains bottlenecked by the time and cost of building specialized assessment datasets. We introduce KNIGHT, an LLM-based, knowledge-graph-driven framework for generating multiple-choice question (MCQ) datasets from external sources. KNIGHT constructs a topic-specific knowledge graph, a structured and parsimonious summary of entities and relations, that can be reused to generate instructor-controlled difficulty levels, including multi-hop questions, without repeatedly re-feeding the full source text. This knowledge graph acts as a compressed, reusable state, making question generation a cheap read over the graph. We instantiate KNIGHT on Wikipedia/Wikidata while keeping the framework domain- and ontology-agnostic. As a case study, KNIGHT produces six MCQ datasets in History, Biology, and Mathematics. We evaluate quality on five criteria: fluency, unambiguity (single correct answer), topic relevance, option uniqueness, and answerability given the provided sources (as a proxy for hallucination). Results show that KNIGHT enables token- and cost-efficient generation from a reusable graph representation, achieves high quality across these criteria, and yields model rankings aligned with MMLU-style benchmarks, while supporting topic-specific and difficulty-controlled evaluation.
title KNIGHT: Knowledge Graph-Driven Multiple-Choice Question Generation with Adaptive Hardness Calibration
topic Computation and Language
Artificial Intelligence
Information Retrieval
68T50, 68T05, 68P20
I.2.7; I.2.4; H.3.3
url https://arxiv.org/abs/2602.20135