An Agentic Evaluation Architecture for Historical Bias Detection in Educational Textbooks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stefan, Gabriel, Dumitran, Adrian-Marius
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911578284097536
author Stefan, Gabriel
Dumitran, Adrian-Marius
author_facet Stefan, Gabriel
Dumitran, Adrian-Marius
contents History textbooks often contain implicit biases, nationalist framing, and selective omissions that are difficult to audit at scale. We propose an agentic evaluation architecture comprising a multimodal screening agent, a heterogeneous jury of five evaluative agents, and a meta-agent for verdict synthesis and human escalation. A central contribution is a Source Attribution Protocol that distinguishes textbook narrative from quoted historical sources, preventing the misattribution that causes systematic false positives in single-model evaluators. In an empirical study on Romanian upper-secondary history textbooks, 83.3\% of 270 screened excerpts were classified as pedagogically acceptable (mean severity 2.9/7), versus 5.4/7 under a zero-shot baseline, demonstrating that agentic deliberation mitigates over-penalization. In a blind human evaluation (18 evaluators, 54 comparisons), the Independent Deliberation configuration was preferred in 64.8\% of cases over both a heuristic variant and the zero-shot baseline. At approximately \$2 per textbook, these results position agentic evaluation architectures as economically viable decision-support tools for educational governance.
format Preprint
id arxiv_https___arxiv_org_abs_2604_07883
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle An Agentic Evaluation Architecture for Historical Bias Detection in Educational Textbooks
Stefan, Gabriel
Dumitran, Adrian-Marius
Artificial Intelligence
Computation and Language
Computers and Society
Multiagent Systems
History textbooks often contain implicit biases, nationalist framing, and selective omissions that are difficult to audit at scale. We propose an agentic evaluation architecture comprising a multimodal screening agent, a heterogeneous jury of five evaluative agents, and a meta-agent for verdict synthesis and human escalation. A central contribution is a Source Attribution Protocol that distinguishes textbook narrative from quoted historical sources, preventing the misattribution that causes systematic false positives in single-model evaluators. In an empirical study on Romanian upper-secondary history textbooks, 83.3\% of 270 screened excerpts were classified as pedagogically acceptable (mean severity 2.9/7), versus 5.4/7 under a zero-shot baseline, demonstrating that agentic deliberation mitigates over-penalization. In a blind human evaluation (18 evaluators, 54 comparisons), the Independent Deliberation configuration was preferred in 64.8\% of cases over both a heuristic variant and the zero-shot baseline. At approximately \$2 per textbook, these results position agentic evaluation architectures as economically viable decision-support tools for educational governance.
title An Agentic Evaluation Architecture for Historical Bias Detection in Educational Textbooks
topic Artificial Intelligence
Computation and Language
Computers and Society
Multiagent Systems
url https://arxiv.org/abs/2604.07883