Abstract
<title>Abstract</title> <p>To address heterogeneous image, creation-statement, and revision-record data in junior high school art assessment, as well as the weak evidence linkage of generative AI feedback, a computer-assisted multimodal feedback evaluation model is developed. The system processes artwork through four computational stages: standard encoding, multimodal recognition, feedback control, and evidence tracking. The recognition module normalizes each input image to 512 × 512 pixels and encodes it into a 768-dimensional visual feature vector, while the text encoder transforms creation statements of up to 300 characters into semantic feature vectors. Cosine similarity is used only to establish initial image–text semantic correspondences rather than to make the final assessment decision. The system subsequently performs rubric-conditioned evidence routing by assigning each candidate correspondence to a specific rubric dimension, evidence region, confidence value, and review state before feedback generation.The feedback module combines assessment deviation, recognition uncertainty, revision urgency, and a teacher-review coefficient to produce tiered revision suggestions. The version-control module uses anonymous identifiers, version numbers, SHA-256 chained verification, role-based permissions, and operation logs to preserve evidence integrity. An eight-week quasi-experiment involving approximately 240 students from six classes generated about 720 artwork versions, 1,200 AI feedback records, and 1,440 teacher blind ratings. After teacher review, the ICC increased from 0.812 to 0.924 and the MAE decreased from 6.84 to 3.27, indicating that multimodal processing, feedback control, and human review can provide a stable computer-assisted evaluation pipeline for subsequent educational-effect analysis. CCS CONCEPTS: Applied computing~Education~Computer-assisted instruction</p>