Abstract
<title>Abstract</title> <p>Background: Marker genes are interpretable summaries of single-cell clusters,but evaluating semantic representations derived from them is vulnerable tolanguage-model leakage, privileged candidate information, and post hoc methodselection. Results: We introduce a staged evaluation framework that preservesnegative findings, closes LLM and deterministic-development branches underexplicit gates, and tests frozen representations on external datasets registeredbefore performance evaluation. DeepSeek free-text reasoning, constrained ontol-ogy selection, and semantic recovery did not demonstrate model-specific valueafter leakage and candidate-universe controls. Four frozen deterministic GeneOntology representations were then evaluated on three independent externaldatasets. All four improved adjusted Rand index over gene-name embeddingson each external dataset. MSOA Top-12 achieved the largest mean Delta ARI(+0.1542), while the strongest representation differed across biological settings.Gains persisted across seven hierarchical settings and 50-seed K-means analyses. Conclusions: The signal that persisted under frozen external validationcame from explicit structured biological knowledge evaluated through a lockedprotocol, not unconstrained generated interpretation or post hoc tuning.</p>