Deprecated: Function curl_close() is deprecated since 8.5, as it has no effect since PHP 8.0 in /home/u483256323/domains/poorvam.com/public_html/subdomains/pore/includes/api.php on line 184
Abstract
<p>Multimodal large language models (MLLMs) are increasingly used to evaluate educational materials, but a holistic rating says little about whether a model responds to a particular instructional-design construct. CFES-P24 is a theory-grounded counterfactual benchmark for auditing educational slides. Its registry specifies 504 planned pairs derived from 24 author-created, three-slide micro-lessons in six disciplines. Of these, 432 target pairs instantiate six multimedia learning principles at three parameterized levels, and 72 are visual-equivalence sham controls. This paper reports the benchmark design and a pilot validation rather than the full 504-pair evaluation. Twenty-one pairs had been generated for this version. They passed 100/100 transformation-rule tests, 21/21 inverse-restoration checks, and 21/21 independent disk-level quality checks. A frozen candidate gate then tested five pairs once with each of two MLLMs, for 10 calls in total. Every response satisfied the structured-output schema on the first attempt. On the eight target calls, observable operation, principle, minimal repair, and evidence-anchor matches were each 8/8; strict direction selection was 6/8, and severity exact match was 0/8. Both sham calls were correctly classified as having no material difference. The preregistered all-or-none gate consequently failed one of 14 rules, so 16 holdback calls were not released. These results suggest that construct recognition, comparative judgment, and severity calibration should be evaluated separately. They also show the importance of distinguishing a registered design from completed evidence when MLLMs are used to audit educational materials.</p>