Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Educational AI systems may receive programme-order information only through sequences of resource documents. We audit whether a small character-level language model can predict a held-out transition between explicitly named curriculummetadata categories whose endpoints, but not target adjacency, were represented during training. Study 1 used a controlled 3 × 4 synthetic subject–topic grid. Repeated factorised metadata transitions improved held-out exact accuracy over a content-matched high-entropy shuffle by 0.425 (95% crossed-bootstrap confidence interval [0.297, 0.545]) across six corpus seeds and five model seeds. Additional deterministic non-factorised paths matched cell-successor entropy, observed-edge count, transition exposure, and text generation; the factorised-minus-control difference was 0.535 [0.385, 0.677]. Length-normalised scoring and free generation gave similar contrasts. No incremental benefit of synthetic lesson bodies was detected when explicit metadata were present, and the result survived an audited counterfactual in which factual and transformed edges were absent from training. Study 2 applied the audit to 383 units from 18 Oak National Academy subject–year programmes. An initially fixed 36-edge study gave an official-minus-shuffle difference of 0.022 [-0.067, 0.128]. A non-confirmatory robustness analysis jointly varying four held-out manifests and four corpus variants produced a pooled difference of 0.365 [0.146, 0.535], but effects were heterogeneous and exact next-unit accuracy remained near zero. Because subject prediction was perfect in both conditions, the robustness-analysis contrast was driven by order-given-subject prediction. As an illustrative design check, a researcher-specified subject/order factorisation recovered every audited official-order metadata target. The evidence supports multi-manifest held-out-edge auditing and identifies the positive result as explicit-metadata factorisation rather than semantic understanding of lesson prose.</p>

Show More

Keywords

heldout study metadata difference audit

Related Articles

PORE

About

Connect